This work combines document parsing, static analysis, and Tree of Thoughts reasoning to make smart contract repair more accurate and practical and overcomes the limitations of linear reasoning.
Abstract
Smart contracts power blockchain applications such as DeFi and NFTs. However, once deployed, they cannot be modified. Even minor bugs can result in significant financial losses. Current AI-based repair methods rely on linear reasoning, which leads to the accumulation of errors and unreliable patches. Our method combines document parsing, static analysis, and Tree of Thoughts reasoning. We first convert audit reports into structured data. Then we use Slither to locate the exact vulnerable code. Our three-step framework explores multiple repair paths simultaneously, evaluates options, and eliminates poor choices. Finally, we verify patches through compilation and manual checks. We test our method on 50 real vulnerabilities from Code4Rena. Our method achieves a 62% single success rate and an 84% top-3 success rate, outperforming ContractTinker by 12 and 6 percentage points, respectively. We also increase the proportion of fully effective patches to 44%, while reducing defective patches from 38% to 22% and invalid patches from 10% to 4%. This approach overcomes the limitations of linear reasoning and makes smart contract repair more accurate and practical.
An executed exploit is reported showing a specification-level authorization defect that produced no high or medium impact finding, its consequence for repair metrics, where it biases both transition counts upward and can confound comparison between methods producing differently sized .patches.
Staley Ian· International Journal of Inn...· 0 citations
Under the single-shot, raw-bytecode-only protocol, current LLMs are not reliable standalone forensic tools, and their robustness has not been systematically tested against contracts adversarially designed to mislead analysis.
Large language models (LLMs) are increasingly applied to the automated repair of C/C++ security vulnerabilities, and compile rate is a commonly reported proxy for progress: whether the generated patch compiles. We argue that compile rate is a scientifically unreliable metric for single-function vulnerability repair, an...
Om Nepal, Sushant Aryal, Oluseyi Olukola et al.· 0 citations
The results show that access to repository history is insufficient: concealment becomes most effective when benign and malicious changes jointly occupy the auditor's active review context or when the stated purpose plausibly accounts for the attack-bearing diff.
Splitting, a black-box wrapper that copies a fuzzer's queue state after a bug trigger and continues from that state in multiple branches, directing more effort toward the discovered region, provides a practical way to measure and improve fuzzers.
Zi-Rui Liu, Meng-Fan Xu, Juan Zhai et al.· 0 citations
Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model"actor"is inspected by a"monitor"(often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's context can steer it to...
Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.