Skip to content
Preprint

ExeCRE: Execution-Consistency Guided Reliability Estimation for Self-Correcting Code Generation

Aug 2026 · 0 citations · 51 references
Computer Science

TL;DR

ExeCRE is an Execution-Consistency guided code Reliability Estimation framework that estimates code reliability by statistically analyzing consistency patterns in execution outputs over a large number of randomly generated inputs, and applies the Dawid-Skene model to infer latent code reliability.

Abstract

Large language models (LLMs) have made notable progress in code generation, but they still struggle on challenging tasks that require sophisticated algorithms or complex implementations. Recent methods increasingly use code execution as feedback, especially in self-correction pipelines that construct verification signals from generated code. However, these pipelines often depend on supervision signals whose reliability is unknown, which can introduce misleading feedback, unnecessary revisions, and incorrect final answers. To address this issue, we propose ExeCRE, an Execution-Consistency guided code Reliability Estimation framework. Instead of judging candidate code by tests or LLM feedback, ExeCRE estimates code reliability by statistically analyzing consistency patterns in execution outputs over a large number of randomly generated inputs. It collects execution outputs over generated inputs, projects them into consistency signals, and applies the Dawid-Skene model to infer latent code reliability. We integrate ExeCRE into self-correction for code generation. Experiments show that ExeCRE consistently improves both effectiveness and stability, while substantially reducing misleading correction signals. Under GPT-5.2 on LiveCodeBench, the average number of misleading feedback cases on already correct code drops from 113.2 with a representative self-correction baseline to 14.0 with ExeCRE. As an additional study, we apply the same reliability estimation strategy to code-based mathematical reasoning and observe similar benefits. These results suggest that ExeCRE enables more reliable use of generated code in execution-based pipelines.

View source

Similar papers

#machine learning Preprint Sep 2026

Introspective Uncertainty Estimation for LLM-Based Code Generation

The findings suggest that hidden states are a robust and informative resource for estimating functional code correctness, supporting a two-stage workflow that combines response-level risk screening with targeted line-level prioritization.

T. Klassert · 0 citations
Review Sep 2026

ExecuCritic: Calibrated Critic Shaping for Code Generation with Verifiable Rewards

Execution feedback is a useful supervision signal for code models because unit tests are objective and directly measure program correctness. Its weakness is that an entire program is often reduced to one pass or fail bit, leaving RLVR to solve a difficult credit assignment problem. At the same time, coding systems ofte...

Jun Cao, Ying-Jie He · 0 citations
#software testing Review Sep 2026

Debugging Functionality-Twisting Translations by LLMs via Differential Testing with Bayesian Prior

Code translation, as a challenging and fundamental task, is increasingly relying on large language models (LLMs). However, LLMs often give seemingly plausible but fallacious translations, misleading and even deceptive to debugging developers. We propose tHinter, an automated approach that frames translation error local...

Shengnan Wu, Xin-Yu Sun, Xin Wang et al. · 0 citations
Preprint Aug 2026

Post-Hoc Attention Steering of Large Language Models for Robust Code Understanding under Obfuscation

This work proposes CodeSteer, a novel attention steering approach that reallocates model attention toward semantically relevant program elements, including backward slices for output prediction and control-flow paths for execution reasoning in large language models.

Xiao-Kai Rong, Aashish Yadavally, Tien N. Nguyen · 0 citations
Preprint Aug 2026

Function-Level Execution Feedback for Code Preference Optimization

STEP-KTODER is proposed, a framework for code preference optimization that defines steps as module-level functions in decomposed multi-function programs and assigns binary correctness labels via automatically generated unit tests and shows that execution-based labels are essential.

Idris Nechnech, Sehwan Kim, Jimin Seo et al. · 0 citations
#natural language process... Preprint Sep 2026

ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation

Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search framework. It organize...

Hui-Fei Wang, Xin-Yi Huang, Yi-Heng Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.