How Benchmarks Mis-Score Computer-Use Agents
A three-tier diagnostic taxonomy shows that verification/feedback and planning failures dominate execution/grounding errors, while a single scalar success rate can not explain, and connects these findings to newer long-horizon CUA benchmarks and derive stage-specific design rules for CUA evaluation.