Aug 2026· International Conference on Multimedia Analysis and Pattern Recognition· pp. 388-393· 0 citations· 20 references
Abstract
Once deployed, smart contract bytecode is final and cannot be patched. Any vulnerability that ships with the contract becomes a permanent attack surface. Existing tools have clear gaps. Static analyzers such as Slither flag aggressively and tend to produce many false positives, while large language models (LLMs) reason fluently but lack audit-specific knowledge. Retrieval-Augmented Generation (RAG) is meant to bridge the two. However, most current pipelines push every retrieved passage into the prompt without any quality check, and when retrieval misses, the injected noise actually makes the output worse instead of helping. We propose a detection pipeline with a retrieval quality gate. Abstract syntax tree (AST) parsing and Slither analysis run in parallel with a two-stage retriever over a 21,032-entry knowledge base built from 208 professional audit reports. A Corrective RAG (CRAG) gate inspects the reranker score and decides whether retrieved evidence is admitted into the prompt at all. On top of that, a Prompt Tiering layer adds audit-derived rules into the Gemini 2.5 Pro context only when CRAG says retrieval is trustworthy. We evaluate on SmartBugs-Curated paired with GPTScan Top200. The full pipeline reaches 80.3% F1 and brings the LLM-only baseline’s false positive rate (FPR) from 44.1% down to 21.2%, which is a 52% relative reduction, while keeping recall at 100%.
A failure analysis of the representation layer underlying GNN-based smart contract vulnerability detectors finds one confirmed case of misclassification caused directly by a representation-layer failure; the prevalence of such failures in real-world contract populations remains an open empirical question.
Birindwa Prisca Hondi, Chinoso Philip Nwishienyi, Charity Wanja Mwaura et al.· 0 citations
Automated feature engineering with large language models (LLMs) can produce semantically meaningful features for tabular data, yet existing methods lack structured domain knowledge, rigorous verification, and explainable provenance. We propose KnowFeat, a knowledge-guided feature engineering framework that organizes do...
Chengsong You, Wangyue Li, Wei-Qiao Que et al.· 0 citations
Large language models (LLMs) embedded in enterprise workflows cannot structurally distinguish legitimate instructions from adversarial ones in the same token stream, making prompt injection OWASP's top LLM risk for two consecutive editions a persistent threat across direct and indirect vectors. This paper presents Prom...
Fatimah Alhamzawi· Al-Noor Journal of Engineeri...· 0 citations
This work recasts vulnerability discovery as an input-prediction task with a closed, deterministic ground truth, and decomposes discovery into three task modes over 22 real-world C/C++ programs spanning 15 domains, finding constraint inference, not navigation, is the dominant bottleneck.
Yuan-Xiang Shi, Jia-Yi Lin, Xuan-Yong Lin et al.· 0 citations
This work evaluates Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, and builds JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code...
Arshak Rezvani, Sasha Behrouzi, Ahmad Sadeghi· 0 citations
Vulnerability-detection benchmarks score the verdict an agent reaches, not the evidence it gathered. A model that recalls a CVE from pretraining therefore scores the same as one that traced the data flow. We study a task where this difference matters, deciding whether a commit introduces a vulnerability. Instead of sco...
Yi-Kun Li, Jin-Feng Jiang, Yuheng Yieh et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.