Preprint
Jul 2026
There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items
This work treats the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items, and introduces the fragility grid, which is a check a leaderboard can run before it reports an order.
V. R. Parupudi
· 0 citations