Skip to content

Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models

Sep 2026 · 0 citations · 32 references
Computer Science

TL;DR

This work uses compression rate to evaluate how well models predict new text, and results yield three main findings: compression performance follows a consistent scaling trend with model size, and lower compression rates are strongly associated with higher zero-shot MMLU accuracy.

Abstract

Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results. Reliable evaluation is particularly challenging for base models, whose limited instruction-following ability complicates task-based assessment. We introduce Uncheatable Eval, a dynamic benchmark that regularly collects newly published text to evaluate base language models and reduce the risk of data contamination. Drawing on the relationship between a model's predictive ability and its ability to compress data losslessly, we use compression rate to evaluate how well models predict new text. We evaluate 80 models across 14 text categories, study how compression changes with context length, and examine the correlation between compression rate and zero-shot MMLU accuracy. Our results yield three main findings: (1) compression performance follows a consistent scaling trend with model size; (2) attention-based, hybrid, and recurrent models differ in how their compression performance changes as more context becomes available; and (3) lower compression rates are strongly associated with higher zero-shot MMLU accuracy. Code is available at https://github.com/Jellyfish042/uncheatable_eval.

View source

Similar papers

Preprint Aug 2026

TokEval: A Tokenizer Evaluation Suite

TokEval is introduced, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics.

Clara Meister · 1 citation
#machine learning Preprint Aug 2026

TokEval: A Tokenizer Evaluation Suite

Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framewor...

Clara Meister · 0 citations
#natural language process... Preprint Sep 2026

Linger and Lose: Knowledge Collapse in Low-Bit Language Models

Training language models with ternary weights is commonly judged by loss and downstream accuracy, which record only a modest cost relative to full precision. We show that these metrics can conceal a much larger failure. We instead measure knowledge capacity, the factual bits stored per parameter, on synthetic biographi...

Prashanna Mani Paudel, S. Sheshappanavar · 0 citations
Preprint Aug 2026

LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression

LowRankArena is presented, a standardized evaluation platform for SVD-based LLM compression that unifies task versions, uniform-precision compression budgets, comparison regimes, and inference measurements, and provides a reproducible pipeline with over 3 TiB released compressed checkpoints.

Zi-Shan Shao, Li-Xun Zhang, Kang-Ning Cui et al. · 0 citations
#machine learning Preprint Sep 2026

Cartridges++: KV Cache Compression without Off-Context Derailment

Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all, ahead of inference t...

Sonia Laguna, João Monteiro, Marco Cuturi et al. · 1 citation
#artificial intelligence Preprint Sep 2026

How Divergence Becomes Decision Flips in Compressed Language Models

Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on the dense model's outputs needs to know how many of its decisions changed. We show that total variation, not KL, answers this directly. Across 802 compressed and perturbed...

Beatriz Almeida Felício · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.