Finding the Bugs Users Would Find: From Sapienz to Autonomous Agents (Keynote)
Abstract
This keynote accompanies the ISSTA 2026 Impact Paper Award for the paper “Sapienz: Multi-objective Automated Testing for Android Applications”, first presented in 2016. Sapienz recast user-interface test generation as a multi-objective search, raising coverage and fault revelation while minimising the length of each failing sequence, so that every reported failure came with a short, replayable witness. Applied to the 1,000 most popular Google Play apps, it revealed 558 previously unknown crashes. The guiding principle: automated testing earns its place only when its results are actionable. What followed was an unusually rapid move from prototype to industrial scale: about three months from publication to the founding of the startup Majicke (acquired by Facebook), and Facebook-scale deployment in 17 months, into continuous integration at Meta, where Sapienz has tested changes, as they are committed to the repository, for the Facebook, Messenger, Instagram, and WhatsApp apps ever since. Today it runs around 2.5 million tests a day across these apps, surfaces roughly 2,000 unique crashes a week, and helps safeguard around 200 app releases a month, thereby ensuring high product reliability and quality. Its reach quickly outgrew crash-finding. Because it generates realistic yet privacy-safe user interactions, Sapienz became the app-exploration engine for dynamic taint analysis at WhatsApp; within the PrivacyCAT system it shifts privacy protection left, helping detect the majority of in-scope privacy incidents before release and safeguarding the privacy of more than two billion users. The talk closes on what Sapienz could not do: judge whether observed behaviour is correct (the oracle problem). Large language model agents can now explore an app and reason about whether behaviour matches what a user should expect, extending automated testing from reliability to functional correctness and, increasingly, to automated repair. Sapienz’s core ideas, multi-objective exploration and actionable fault revelation, carry over directly. A decade on, that founding principle still holds true.