This work uses compression rate to evaluate how well models predict new text, and results yield three main findings: compression performance follows a consistent scaling trend with model size, and lower compression rates are strongly associated with higher zero-shot MMLU accuracy.
A taxonomy of reasoning enhancement techniques is proposed, categorized into training-time strategies (e.g., supervised fine-tuning, reinforcement learning) and test-time mechanisms (e.g., prompt engineering, multi-agent systems), and outlining future directions toward building efficient, robust, and sociotechnically r...
Zi-Zhan Ma, Wen-Xuan Wang, Meidan Ding et al.· 16 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.