Preprint
Aug 2026
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
This work systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains and establishes StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.
Liya Zhu, Xin Ma, Tao Liu et al.
· 0 citations