Skip to content
Preprint

Mitigating Error Propagation in Chain-of-Thought: A Tree-of-Thought Framework for Smart Contract Repair

Aug 2026 · 0 citations · 70 references
Computer Science

TL;DR

This work combines document parsing, static analysis, and Tree of Thoughts reasoning to make smart contract repair more accurate and practical and overcomes the limitations of linear reasoning.

Abstract

Smart contracts power blockchain applications such as DeFi and NFTs. However, once deployed, they cannot be modified. Even minor bugs can result in significant financial losses. Current AI-based repair methods rely on linear reasoning, which leads to the accumulation of errors and unreliable patches. Our method combines document parsing, static analysis, and Tree of Thoughts reasoning. We first convert audit reports into structured data. Then we use Slither to locate the exact vulnerable code. Our three-step framework explores multiple repair paths simultaneously, evaluates options, and eliminates poor choices. Finally, we verify patches through compilation and manual checks. We test our method on 50 real vulnerabilities from Code4Rena. Our method achieves a 62% single success rate and an 84% top-3 success rate, outperforming ContractTinker by 12 and 6 percentage points, respectively. We also increase the proportion of fully effective patches to 44%, while reducing defective patches from 38% to 22% and invalid patches from 10% to 4%. This approach overcomes the limitations of linear reasoning and makes smart contract repair more accurate and practical.

View source

Similar papers

Open access Aug 2026

Three Measurement Hazards in Analyzer-in-the-Loop Repair of Smart Contracts

An executed exploit is reported showing a specification-level authorization defect that produced no high or medium impact finding, its consequence for repair metrics, where it biases both transition counts upward and can confound comparison between methods producing differently sized .patches.

Staley Ian · 0 citations
#artificial intelligence Preprint Sep 2026

Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

Large language models (LLMs) are increasingly applied to the automated repair of C/C++ security vulnerabilities, and compile rate is a commonly reported proxy for progress: whether the generated patch compiles. We argue that compile rate is a scientifically unreliable metric for single-function vulnerability repair, an...

Om Nepal, Sushant Aryal, Oluseyi Olukola et al. · 0 citations
Review Aug 2026

PRWeaver: Evaluating LLM-Based Code Auditors against Long-Horizon Malicious Pull Requests

The results show that access to repository history is insufficient: concealment becomes most effective when benign and malicious changes jointly occupy the auditor's active review context or when the stated purpose plausibly accounts for the attack-bearing diff.

Yuekun Wang, Ming-Fei Cheng, Xiao-Fei Xie · 0 citations
Preprint Sep 2026

Robustness-Aware Evaluation and Enhancement of Mutation-Based Fuzzing for Bug Discovery

Splitting, a black-box wrapper that copies a fuzzer's queue state after a bug trigger and continues from that state in multiple branches, directing more effort toward the discovered region, provides a practical way to measure and improve fuzzers.

Zi-Rui Liu, Meng-Fan Xu, Juan Zhai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model"actor"is inspected by a"monitor"(often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's context can steer it to...

Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.