Skip to content
Open access

When Does Context Help? A Controlled Study ofLLM-Based Bug Fixing

Sep 2026 · EAI Endorsed Transactions on Internet of Things · 0 citations · 26 references

Abstract

Large language models (LLMs) have shown promise for automated program repair, but it remains unclear which debugging signals are most useful and when additional context becomes distracting, costly, or ineffective. We present a controlled empirical study of LLM-based bug fixing on FIXEVAL, comparing three model families–GPT, Claude, and DeepSeek–under five prompt settings: code-only repair, problem description, failed test cases, passed and failed test cases, and all available context. We further investigate whether structured execution feedback improves repair in a second attempt after an initial patch fails. Across 400 benchmark instances spanning Wrong Answer, Runtime Error, Time Limit Exceeded, and Memory Limit Exceeded bugs, we evaluate repair accuracy, response time, computation cost and second-attempt recovery. The results show that contextual information generally improves first-attempt accuracy, but its benefit depends strongly on bug type and model family. Full context is often most effective for Wrong Answer, Runtime Error, and Time Limit Exceeded bugs, whereas Memory Limit Exceeded bugs remain difficult across settings, suggesting that resource-limit failures often require deeper algorithmic redesign. Passed tests provide limited benefit unless paired with stronger semantic or failing-test signals. Structured failure feedback improves second-attempt repair most when it exposes concrete behavioral mismatches or runtime failures, but it is less effective for coarse resource-limit signals. These findings clarify when context helps LLM-based repair and provide practical guidance for designing cost-aware, feedback-driven repair pipelines.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.