Skip to content
Open access

The genetic architecture of human programmed stop codon readthrough

Jul 2026 · bioRxiv · 0 citations
Biology

TL;DR

This study provides comprehensive quantitative maps of the sequence determinants of human programmed readthrough and suggests that three examples of programmed readthrough are located on distinct local fitness peaks, each defined by different upstream and downstream architectures built around a shared CUAG core motif.

Abstract

Programmed translational readthrough produces C-terminally extended protein isoforms via decoding of stop codons by near-cognate tRNAs. Human genes experimentally validated as readthrough targets share a CUAG motif downstream of a UGA stop codon. However, the full sequence determinants of readthrough efficiency, how they combine, and how generalisable they are across genes remain largely unexplored. Here we use deep mutational scanning to quantify ∼1,400 sequence variants for each of the three examples of human readthrough in the genes AQP4, MAPK10 and OPRK1. In addition to the core CUAG motif, mutations that modulate readthrough elements extend up to +27 nucleotides downstream of the stop codon and across six codons (18 nucleotides) upstream. For the downstream sequence, an additive model with a sigmoidal global epistasis function captures most of the within-gene readthrough variance for double mutants (R²=0.84-0.96), with additional contributions from a small number of strong pairwise interactions. Mutational effects nonetheless generalise poorly between genes: only the immediate -3 to +4 nucleotide window shows consistent behaviour, while mutations in more distal positions have context-dependent effects. Combinatorial assembly of sequence blocks from different genes into chimeras reveals strong interactions (epistasis) between sequences upstream and downstream of the stop codon. This study provides comprehensive quantitative maps of the sequence determinants of human programmed readthrough and suggests that three examples of programmed readthrough are located on distinct local fitness peaks, each defined by different upstream and downstream architectures built around a shared CUAG core motif.

Read PDF

Similar papers

Review Open access Aug 2026

RNA Cis-Elements Involved in Animal Virus Stop Codon Readthrough: Stop Codon Context and Downstream RNA Structures

Stop codon readthrough is a noncanonical translation strategy employed by certain RNA viruses, in which a viral termination codon is either decoded by host near-cognate tRNAs or canonically recognized by the class I release factor (RF, eRF1 in eukaryotes). Ribosomal A-site competition between near-cognate tRNAs and eRF can shift decoding toward near-cognate tRNAs, thereby promoting non-canonical decoding events by transiently pausing termination and favoring readthrough. This review focuses on two viral cis-elements that modulate readthrough across four viral genera in which this decoding event has been experimentally validated: (i) primary sequences surrounding the stop codon (stop codon context), and (ii) downstream RNA structures. Effects of stop codon context have been observed more broadly in cellular genes, including nonsense suppression in bacteria, with mechanisms including inefficient RF association or tRNA interactions at adjacent sense codons. In eukaryotic systems, interactions with the ribosomal mRNA entry channel have been suggested. Diverse downstream structures, including gammaretroviral pseudoknots and specific structures in alpha- and coltiviruses, further stimulate readthrough in a location- and structure-sensitive manner. This effect has not been consistently observed in chikungunya and triatoviral structures, suggesting a strong dependence on local sequence and structural context. Compared with the larger number of cellular readthrough occurrences that can be detected at low efficiency by ribosome profiling, viral readthrough in mammalian systems is consistently high (>2%). Understanding the interplay between viral RNA elements and host translational machinery, including potential kinetic trapping at the termination codon, provides insights into this unusual elongation mechanism. These findings may have implications for antiviral strategies targeting these RNA elements.

N. Kamoshita · 1 citation
Open access Aug 2026

Codon optimization depletes stop codons in alternative reading frames of protein-coding nucleic acid therapeutics.

Out-of-frame translation events, arising from ribosomal frameshifting or noncanonical initiation, are an unavoidable feature of translation. Their consequences depend on the distribution of stop codons in alternative reading frames, which determines the permissiveness of those frames: whether out-of-frame translation terminates quickly or generates extended products. Using quantitative dual-fluorescence reporters, we show that these stop codons function as molecular checkpoints that terminate out-of-frame translation. Genome-wide analysis across 10 organisms reveals that natural coding sequences maintain dense stop codon distributions in alternative frames, with a median spacing of approximately 20 amino acids. Codon optimization, the standard method for enhancing translation, systematically depletes this safeguard. Because all three stop codons (UAA, UAG, UGA) begin with uridine, and optimal human codons exclude uridine from third positions, stop codons in the -1 reading frame become structurally impossible in codon-optimized sequences. Analysis of 120 therapeutic sequences, including FDA-approved COVID-19 messenger RNA (mRNA) vaccines, confirms widespread -1 frame stop codon depletion: out-of-frame products average 164 amino acids, sixfold longer than in natural human genes. Strategic restoration of stop codons through synonymous substitutions eliminates detectable out-of-frame products by mass spectrometry while preserving the intended protein. Although such products and immune responses have been detected in COVID-19 mRNA vaccine recipients, there is no evidence they cause clinical harm; nonetheless, our approach offers a simple way to eliminate them through informed sequence design alone, without changes to manufacturing or regulatory frameworks. Our findings establish stop codon distribution as a critical design parameter for protein-coding nucleic acid therapeutics.

Zheling Liu, Zhenguang Ying, Luoan Shen et al. · 0 citations
Open access Aug 2026

MetaTIS: a tool to predict cognate and near-cognate translation initiation sites in human

Abstract Ribosomes typically commence translation at a methionine-encoding AUG codon flanked by a so-called Kozak region, a short nucleic acid motif that serves as an initiation site in humans. Though, the characteristic AUG start codon of an mRNA is not always effective in initiating translation. Near-cognate codons differing from AUG by one nucleotide may also be recognized as start sites. Several types of ribosomal profiling techniques have been developed that elucidate active translation initiation sites (TIS) that enable training of computational models to predict both cognate and near-cognate TIS using mRNA sequence features. Here, a meta-model termed MetaTIS was implemented by combining outputs of genomic and protein language models fine-tuned on Ensembl annotations of transcripts and five different TIS datasets. The model proficiently differentiates between spurious and true TIS in four distinct test sets, for both AUG and non-AUG instances. Most important for translation initiation based on one of the base model outputs was the Kozak sequence context and a region further upstream in the 5′UTR [−12, −10]. MetaTIS is available as a webserver at https://service2.bioinformatik.uni-saarland.de/metatis/, a tool that accurately predicts TIS for AUG and nine near-cognate start codons.

Aram N. Papazian, Volkhard Helms · 0 citations
Open access Sep 2026

Rapid Translation of the Early Coding Region Promotes Premature Transcription Termination, Reducing Protein Expression in Escherichia coli.

Slowly translated codons are overrepresented among the first approximately 30 codons of natural genes across all domains of life, yet the functional basis for this conserved feature remains incompletely understood. Using the Escherichia coli lacZ gene as a model, we previously showed that insertion of fast-translated codons into the early coding region dramatically reduces β-galactosidase production by increasing premature transcription termination and decreasing mRNA stability. Here we show that premature transcription termination events occur several nucleotides downstream of the fast-codon inserts, at positions coinciding with previously characterized Rho-dependent intragenic termination sites in lacZ. To exclude the possibility that the specific amino acid sequence of the inserts-rather than their translation speed-was responsible for termination, we constructed two additional lacZ variants with distinct fast-translated sequences, including one replacing the 30 earliest lacZ codons with their fastest synonymous alternatives. Both variants showed premature transcription termination of comparable magnitude to the original inserts, demonstrating that it is the high translation rate itself which causes premature termination. Analysis of a conservative set of seven high-confidence fast-translated codons across natural highly expressed genes and the synthetic constructs revealed that a run of consecutive fast codons is the feature that most clearly distinguishes the terminating sequences from natural genes. We propose that excessively rapid synthesis of the N-terminal part of a polypeptide may impair its proper entry into the ribosomal exit tunnel, thereby disrupting transcription-translation coupling and exposing downstream mRNA to Rho-dependent termination.

S. Pedersen, Alberte Honoré Jepsen, Bertil Gummesson et al. · 0 citations
Open access Aug 2026

Sequence-intrinsic barriers define a class of translation-restricted uORFs in the human genome

This work systematically quantified the intracellular accumulation of 111 evolutionarily conserved human uORFps and identified a conserved class of human uORFs that encode intrinsic barriers to productive translation and provide a rigorous framework for understanding how noncanonical coding sequences shape the human proteome.

Hinata Hashimoto, Taichi Akase, H. Kurasawa et al. · 0 citations
Open access Aug 2026

Pangenome discovery and characterization of human protein-coding duplicated genes

The pangenome provides unparalleled specificity to understand genetic variation in SD genes allowing us to distinguish functional genes from pseudogenes and highlighting potential gene innovations that arose most recently in human evolution.

Luyao Ren, DongAhn Yoo, Katarina Vlajic et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.