Skip to content
Open access

A leakage-controlled benchmark shows apparent codon-language-model advantages in synonymous-variant prediction are evaluation artifacts

Aug 2026 · bioRxiv · 0 citations · 32 references
Biology

Abstract

Synonymous codon choices shape mRNA stability, translation, and folding, and codon language models (cLMs) are increasingly reported to read this biology from sequence. However, when a true signal is thin relative to a confounding one, standard evaluation protocols can manufacture the reported gain rather than measure it—and we show this is what has happened for cLMs on synonymous-variant prediction. Under random splits, the codon advantage is large: tokenization gaps of +2.3–14.3 percentage points (pp) and pretrained codon leads of +2.9 pp over the strongest protein model (ESM-1b) and up to +4.9 pp over ESM-2. We find these numbers are properties of the measurement, not the models. A memorization baseline outscores every neural model; the advantage collapses under gene-held-out evaluation; the sole surviving residual dissolves into six defensible probe defaults; and the synonym-randomization drop that appeared to confirm true signal is itself variance under pooled analysis (0.3 pp, p = 0.49). No advantage survives leakage-controlled evaluation with pooled statistics. We release CodonBench, a leakage-controlled benchmark with an emergent audit cascade, and characterize how artifacts accumulate at every pipeline step. A thin signal (I(σ; Y |A) ≈ 0.04 bits) may exist but is not reliably detectable at current sample sizes; we specify what detecting it would require.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.