Preprint
Sep 2026
The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean
The results suggest that benchmark scores should be interpreted together with their scaffolding level, scoring criterion, and reliability profile, providing a practical framework for more valid evaluation of LLM agents.
Yong-Hong Zhang, Shadi Motaali, Vu Phong Dinh et al.
· 2 citations