2026· IEEE Transactions on Software Engineering· pp. 1-13· 0 citations· 83 references
TL;DR
A controlled human-subjects study examining the effects of physical exercise and AI-generated code summaries with varying correctness on developers’ summarization and bug detection performance concludes that AI assistance is generally useful for code summarization, but incorrect AI assistance substantially degrades bug detection performance.
Abstract
—Code summarization supports program comprehension and defect detection, and developers increasingly use both human-centered and AI-based interventions when performing this task. We present a controlled human-subjects study ( N = 47 ) examining the effects of physical exercise and AI-generated code summaries with varying correctness on developers’ summarization and bug detection performance. Participants summarized GitHub code under different intervention conditions, with outcomes evaluated along multiple dimensions, including accuracy, completeness, conciseness, readability, defect detection, and response time. We analyzed the data using mixed-effects models to account for repeated measures across participants and code artifacts. Surprisingly, under the exercise intervention studied, we did not observe consistent benefits for summarization quality. AI assistance is generally useful for code summarization, but incorrect AI assistance substantially degrades bug detection performance. These findings provide empirical evidence on the nuanced benefits and risks of human-and tool-based interventions in code summarization and bug detection.
AlHAZEN and AVICENNA show that effective debugging tools have tradeoffs between accuracy and interpretability to support developers’ decision-making in increasingly complex software environments.
C. Lazik, Martin Eberlein, Aaron Ziglowski et al.· Message Understanding Confer...· 0 citations
Recently, Developers have been relying on AI tools to support them in their daily work by generating code. While the use of large language model-based AI tools has improved productivity, the quality of the generated code wasn't always optimal. In a lot of cases, the code includes design issues known as code smells, whi...
Y. Younes, Yousef Elsheikh· IEEE Jordan Conference on Ap...· 0 citations
Evaluated SWIRL, which supports interpreting tool-generated warnings through interactive, customized summarization, suggests that SWIRL's active learning-based summarization can enhance the sensemaking process of tool-generated warnings.
This study compares the structural quality of code produced by three widely adopted vibe coding tools --- Lovable, v0, and Replit --- starting from a single generation prompt and suggests that choosing between vibe coding tools involves structural trade-offs that go beyond perceived productivity.
These findings validate cognitive theory for explainable, actionable, and interpretable safety-critical defect prediction, laying empirical groundwork to evaluate analogous issues in LLM-generated code through the behavioral study of AI.
Carlos Andrés Ramírez Cataño, Makoto Itoh· International Conference on...· 0 citations
This work introduces RepoProbe, a novel benchmark for evaluating repository-level code understanding through open-ended Q&A using GitHub Discussions, which focuses on open-ended architectural inquiries rather than defect reporting and proposes a Checklist-Based Verification Protocol that decomposes answers into atomic,...
Yue Yang, Alyssa Wu, Ji Luo et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.