Automated Program Repair for UI-Centric Android Bugs: How Far Are We?
Abstract
Substantial research effort has been devoted to developing techniques for automated program repair (APR) that suggest patches for localized buggy code -- and more recent techniques have begun to leverage the capabilities of code-centric large language models (LLMs). However, the scope and diversity of bugs to which these techniques have historically been applied are limited. In particular, the research community currently lacks a comprehensive understanding of the performance of APR techniques on bugs that arise in UI-centric programs, such as mobile apps. Bugs in UI centric programs carry with them unique challenges, including (i) the need to reason across interconnected subroutines that connect presentation and program logic, (ii) event-driven programming paradigms, and (iii) the need to reason about program state through cues in the UI. In this paper, we investigate the effectiveness of existing APR techniques when applied to fix bugs in UI-centric programs - specifically Android applications. To explore this phenomenon, we conduct a comprehensive empirical study with five existing program repair techniques (including those that utilize LLMs) on a hybrid dataset including 46 synthetic bugs, generated via MDroid+, an Android-specific mutation tool, and 50 real bugs systematically mined from issue reports of 23 popular Android applications. Our findings illustrate important current limitations in resolving UI-related issues in mobile apps. We synthesize these results to form a taxonomy of the limitations of existing program repair techniques. This taxonomy outlines key limitations and can inform future research efforts in designing automated program repair tools for UI-centric bugs in mobile applications.