RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
∗ YanshiLi, ∗. XueruBai, Shuman Liu et al.
· 0 citations
2 papers indexed here
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
Under cross-benchmark transfer evaluation across twelve models completed by all four readouts, difference-in-means attains the highest model-level mean on ten models, while logistic regression wins the most capability-model cells.