Extensive evaluations across three representative vulnerability cases and ten backdoor attacks, along with sixteen competitive baselines, demonstrate that CodeTracer consistently achieves high forensic accuracy, low false identification rates, and strong robustness against adaptive attacks.
Abstract
Large language models have enabled powerful code completion systems that assist developers by predicting subsequent lines of code. However, these models remain vulnerable to backdoor attacks, where malicious fine-tuning data covertly implants unsafe behaviors. Despite advances in defensive techniques, adaptive and sophisticated backdoor attacks still evade detection and mitigation. We present CodeTracer, a forensic framework that traces malicious code completions back to the backdoor fine-tuning data responsible for them. Operating under realistic post-deployment constraints, CodeTracer relies solely on the fine-tuning corpus and the reported miscompletion event. It extracts a structured behavioral fingerprint from the compromised output, narrows the search to semantically relevant code samples, and employs LLM-based reasoning to attribute unsafe logic to specific backdoor data. Extensive evaluations across three representative vulnerability cases and ten backdoor attacks, along with sixteen competitive baselines, demonstrate that CodeTracer consistently achieves high forensic accuracy, low false identification rates, and strong robustness against adaptive attacks.
Semantically-Equivalent Transformation (SET)-based backdoor attacks are introduced, a new class of attacks that use semantics-preserving low-prevalence code transformations to generate stealthy triggers and are proposed as a framework for constructing and prioritizing such triggers.
Junyao Ye, Zhen Li, Xi Tang et al.· ACM Transactions on Software...· 0 citations
BADERASER is proposed, a novel backdoor defense technique for backdoor elimination in neural code models that introduces code naturalness as an auxiliary constraint and incorporates statistical indicators in trigger inversion to improve the quality of recovered triggers.
Wei Cheng, Yu Zhou, Guang Yang et al.· International Conference on...· 0 citations
LLM customization platforms allow users to build task-specific models for code intelligence tasks by embedding instructions into system prompts, without modifying the underlying model parameters. While these platforms lower the barrier to developing customized LLMs, they also introduce a new attack surface: instruction...
Yuchen Chen, Wei Cheng, Yuan Xiao et al.· 0 citations
Lily is presented, an automated approach that strengthens open-source development and release processes against backdoor injection and achieves high detection accuracy with low false alarm rates, reliably identifies malicious code, resists adversarial attempts, and would have prevented real-world backdoor incidents.
Dimitrios Kokkonis, M. Marcozzi, Stefano Zacchiroli· arXiv.org· 0 citations
This work presents SCRIPTIOC-BENCH, a benchmark for measuring static IOC extraction capability on real-world malicious scripts, and evaluates a broad range of proprietary and open-weight LLMs, showing that IOC recovery without execution remains challenging across model scales.
Hanna Kim, Jian Cui, Minkyoo Song et al.· 0 citations
A dual-stream CodeBERT architecture is presented that addresses keyword bias —by combining pre-trained Transformer representations with a 50-dimensional hand-engineered security feature vector, supported by targeted data augmentation and two-stage adversarial fine-tuning.
Arjun Khurana, Talaya Farasat, Joachim Posegga et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.