Skip to content
Preprint

The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean

Sep 2026 · 2 citations · 49 references
Computer Science Physics

TL;DR

The results suggest that benchmark scores should be interpreted together with their scaffolding level, scoring criterion, and reliability profile, providing a practical framework for more valid evaluation of LLM agents.

Abstract

Agent benchmarks are increasingly used to compare large language models (LLMs) and guide deployment decisions, yet benchmark scores are meaningful only if they measure model capability rather than properties of the evaluation pipeline. We identify a double measurement confound: execution-critical decisions are performed by a fixed scaffold instead of the model, while the scorer evaluates outputs using criteria that may not reflect task correctness. We unify these issues within a measurement-theoretic framework that characterizes when benchmark scores can be interpreted as evidence of model capability, and instantiate it with an audit-and-repair protocol that (i) transfers execution-critical decisions from the scaffold to the model, (ii) replaces shape-based evaluation with seeded ground-truth scoring, and (iii) reports reliability beyond the mean through worst-case and tail-risk metrics. Experiments on ComtradeBench show that the joint intervention transforms a nearly flat leaderboard into a reliability spectrum that distinguishes both average performance and robustness across seeds. Applying the audit to existing benchmarks further shows that scorer validity is benchmark-specific, whereas scaffold ownership is an uncontrolled axis wherever we probed it. Our results suggest that benchmark scores should be interpreted together with their scaffolding level, scoring criterion, and reliability profile, providing a practical framework for more valid evaluation of LLM agents.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

DeltaSelect: Affordable A/B Testing for Coding Agents

Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions. Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice. In a resampling analysis of DeepSWE's published trials, only 19.5% of tasks (22 of 113)...

Nicholas J. Conn · 0 citations
Preprint Aug 2026

Demystifying Agent Skills: Why They Work-Until They Don't

This work designs a contrastive study that combines controlled quantitative experiments with paired trajectory analysis and consolidates observations into a taxonomy of three high-level categories and twelve skill-use modes, showing that skills work when noisy trajectories become procedural anchors that stabilize execu...

Zhi-Yuan Jiang, Fan Huang, Hanwen Xing et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Self-Designed Evaluators and Warm Memory for Long-Horizon Agents

A tool-using language-model agent deployed over a long stream of tasks receives no reward, so it cannot tell whether it succeeded, cannot safely retry, and cannot label the experience it needs to improve. We present SelfSuite, in which the agent's own base model, given only the world's public materials, designs a small...

Saeid Asgari, Emre Kıcıman, Leonardo de Oliveira Nunes et al. · 0 citations
#artificial intelligence Preprint Aug 2026

K-Bench: measuring model performance on real scientific agent requests

K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs is reported, arguing that the informative quantity for scientific agents is not a leaderboard position but the joi...

Aubrey M. Brueckner, D. Patel, Yuhuan He et al. · 2 citations
#machine learning Preprint Aug 2026

Predicting Task Difficulty Without Rollouts

This work shows that AUC can mask poor difficulty estimates, identify token-level entropy as a useful predictive signal, and show how residuals between expected and observed difficulty can expose hidden environment flaws such as contamination and infeasibility.

Stefan Krsteski, Charlotte Meyer · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.