This work presents UpgradeBench, a decision-driven longitudinal benchmark covering four consecutive Qwen releases, one continuation checkpoint, six tasks, and two model sizes, augmented by OLMo checkpoints with known training lineage, and disentangles three core questions: whether a new checkpoint improves fixed-recipe retrained specialist performance, whether specialization assets transfer across versions, and what recovery resources are usable.
This work formalizes evaluator reasoning accountability via three core sources: grounds, norms, and authority, and defines judgment receipts as minimal source replacement sets that reproduce revised verdicts to explain judgment transitions.
Yefei Chen, Wei-Ning Zhang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.