Preprint
Jul 2026
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
E-Bench is introduced, a fully synthetic benchmark with 323 state-changing tasks across three product domains: Honor of Kings, QQ Music, and Tencent Meeting, and it shows that multi-step tool use remains challenging: Pass^3 stays below 60% for the strongest models, and even with code execution in the E-Bench-Code extension, reliability remains below 70%.
Weihuang Zheng, Tianyuan Zou, Eileen Ye et al.
· 1 citation