Skip to content

Analyzing and fixing code generation errors of foundation large language models

Sep 2026 · Empirical Software Engineering · Vol 32 · 0 citations · 56 references

TL;DR

The goal is to understand the code generation errors of foundation LLMs and explore the solution to resolve directly fixable errors, and to design and evaluate the LlmFix fixing method and constructed the LlmErrorEval dataset.

View source

Similar papers

Open access Aug 2026

Improving Bug Detection in LLM-Generated Unit Tests: Revisiting Test-Oracle Reliability Across Modern Large Language Models

This paper presents a formal mathematical model for categorizing the outcome of generated-tests into four classes, a couple of basic metrics: Bug-Revealing Rate (BRR) and Bug-Validating Rate (BVR); and two basic statistical tests to ensure that the results are rigorous.

Zeyad Farooq Lutfi · 0 citations
#natural language process... Preprint Sep 2026

Retrofitting Code Using LLMs to Support Exceptional Behavior

Exception Related Code (ERC), which includes throw statements, conditions (if statements) that guard those throw statements, and try/catch blocks, is an essential component of software systems, allowing developers to detect and handle exceptional states that deviate from the expected program behavior. However, manually...

Ling-Tao Zhong, Jiyang Zhang, Jayanth Srinivasa et al. · 0 citations
#natural language process... Preprint Sep 2026

Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in problem solving or bug fixing independently, but do not explore the relationship between thes...

Alexandru Stefan Stoica, Traian Rebedea, M. Mihăescu · 0 citations

Hunk-Constrained DPO: Segment-Level Optimization for Secure and Correct LLM Code Generation

Hunk-Constrained Direct Preference Optimization is introduced, a training framework that unifies security hardening and functional correction in large language models and demonstrates that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.

Qian-Shuo Huang, Xin Yin, Xin-Rui Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Interpreting and Steering for Safe and Correct Code Generation

DuoSteer is proposed, a double-steering approach that simultaneously applies safety and code-correctness steering to attention heads and outperforms not only other steering variants but also prompting and supervised fine-tuning baselines for inference-time vulnerability reduction.

Hao Yan, Zi-Yu Yao · 0 citations
Open access Sep 2026

Goanna: a novel approach for automated type error debugging

Statically typed languages offer many advantages in software engineering, including bug prevention, enhanced code quality, and reduced maintenance costs. However, these benefits come at the expense of a steep learning curve and a slower development pace. Although known for its expressive and strong type system, Haskell...

Shuai Fu, Tim Dwyer, Peter James Stuckey et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.