Preprint
Jul 2026
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
Under cross-benchmark transfer evaluation across twelve models completed by all four readouts, difference-in-means attains the highest model-level mean on ten models, while logistic regression wins the most capability-model cells.
Yanshi Li, Xue Bai, Shuman Liu et al.
· 0 citations