APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents
Mobile GUI agents can execute tasks from natural-language instructions, but their evaluation remains difficult to make both realistic and reproducible. Existing benchmarks typically trade off these goals: simplified apps lack real-world mobile complexity, whereas live commercial apps introduce uncontrolled variation fr...