Evaluated SWIRL, which supports interpreting tool-generated warnings through interactive, customized summarization, suggests that SWIRL's active learning-based summarization can enhance the sensemaking process of tool-generated warnings.
Abstract
Programmers using bug-finding tools often review their reported warnings one by one. Based on the insight that identifying recurring themes and relationships can enhance the cognitive process of searching for representations of a given problem space (i.e., sensemaking), we propose SWIRL, which supports interpreting tool-generated warnings through interactive, customized summarization. With active feedback, SWIRL derives summary rules for grouping of related warnings on the fly. As users mark warnings as interesting or uninteresting, SWIRL's rule inference algorithm surfaces common characteristics, highlighting structural similarities in containment, subtyping, invoked methods, accessed fields, and expressions. We demonstrate SWIRL on real-world warnings generated from Infer and SpotBugs on two mature Java projects. In a within-subject user study, our participants articulated root causes for similar uninteresting warnings with more confidence when using SWIRL, compared to the baseline that lists individual warnings without customized summary rules. Among participants, we observed significant individual variation in desired grouping, reinforcing the need for individualized sensemaking. The simulation we conducted shows that SWIRL's rule-level feedback expedites sensemaking, requiring only 11.8 interactions on average to align all inferred rules with a simulated user's labels when combined with instance-level feedback, compared to 17.8 interactions when using instance-level feedback alone. Our evaluation suggests that SWIRL's active learning-based summarization can enhance the sensemaking process of tool-generated warnings.
Monitoring online platforms, one of the most challenging aspects of the fact-checking process, remains predominantly manual. Crafting search queries to surface potentially dubious online content is based on guesswork, demanding significant time and effort from fact-checkers. While monitoring tools exist for platforms l...
Prerna Juneja, Dong-Kai Xie, Guang-Yin Ye et al.· Proceedings of the ACM on Hu...· 0 citations
An evaluation paradigm that assesses model responses along three criteria grounded in cooperative response theory is designed, suggesting pragmatic redirection is a fundamentally underdeveloped capability in current LLMs.
Akhila Yerukola, Jena D. Hwang, Ming-Qian Zheng et al.· 1 citation
Comprendia is an Eclipse plugin that integrates structural dependency visualization with LLM-powered code explanation on a shared interactive graph for Java program comprehension. The tool rests on four pillars: (1) a multi-edge-type dependency graph with live search and multiple layouts; (2) LLM explanations grounded...
SmellCC, a Visual Studio Code extension that augments SonarQube with an LLM-based pipeline to automatically detect and refactor Python code smells, provides in-place, one-click remediation for the top-10 most frequent smells, effectively preventing the accumulation of technical debt during development.
Xiaoting Zhang, Yujie Zhang, Zhi-Peng Gao et al.· 0 citations
A lightweight quality-assessment protocol is presented for LLM-generated synthetic training data and applied to 13,579 synthetic user reviews generated from GitHub issues across four open-source Android applications, high-lighting the need for hybrid human-AI verification when synthetic data is used in security-critica...
Ogtay Hasanov, Saad Ezzini· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.