Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models
This work uses compression rate to evaluate how well models predict new text, and results yield three main findings: compression performance follows a consistent scaling trend with model size, and lower compression rates are strongly associated with higher zero-shot MMLU accuracy.
Kai-Feng Tan, Yu-Dong Li, LinLin Shen
· 0 citations