Skip to content
Open access

Recovering vulnerability detection on realistically-distributed code

TL;DR

This thesis rebuilds Real-Vul through a controlled Code Property Graph pipeline, measure and remove a 37% content leak intrinsic to whole-codebase sampling, and train a relational graph neural network with a disciplined class-imbalance recipe: focal loss, class-aware undersampling, and fine-tuning of a GraphCodeBERT node encoder.

Abstract

Prior deep-learning vulnerability detectors report F1 scores (the harmonic mean of precision and recall) as high as 99% on benchmark datasets, yet collapse to the trivial all-safe classifier when evaluated on realistically-distributed code. On Real-Vul (Chakraborty et al., 2024), a corpus that simulates deployment-scale class imbalance of roughly 0.3% positive, every published model studied scores an area under the receiver operating characteristic curve (ROCAUC) near 0.50 and vulnerable-class F1 near zero. This thesis demonstrates that the collapse is not fundamental.We rebuild Real-Vul through a controlled Code Property Graph (CPG) pipeline, measure and remove a 37% content leak intrinsic to whole-codebase sampling, and train a relational graph neural network with a disciplined class-imbalance recipe: focal loss, class-aware undersampling, and fine-tuning of a GraphCodeBERT node encoder. On the deduplicated test set (n = 30,873; 518 vulnerable; 1.68% positive), the model reaches ROC-AUC 0.975 (95% CI [0.967, 0.983]) and vulnerable-class F1 0.773 (95% CI [0.746, 0.799]), where the strongest prior result on the same corpus reached F1 0.46 / AUC 0.82.Three controlled experiments locate the cause and the boundary of the result. An architectural ablation attributes the gain to the training recipe operating on a protocol-matched corpus; swapping the readout architecture leaves the metrics unchanged. A four-corpus transfer matrix shows that the result holds only within a single labeling protocol. A Java imbalancereproduction experiment shows the recipe consistently beating no-handling on the same corpus while reaching a strong detector only where the corpus carries learnable structural signal at scale. The result is a working detector on realistic data, with an account of the conditions under which it works.

Read PDF

Similar papers

2026

LLMKernelBench: Benchmarking Large Language Models on Software Vulnerability Detection in Linux Kernel

Large language models (LLMs) demonstrate strong capabilities in code-related tasks, however their effectiveness in software vulnerability detection (SVD) remains poorly understood due to inadequate evaluation frameworks. Existing benchmarks suffer from training data contamination, isolated function evaluation without c...

Arastoo Zibaeirad, Rodrigo Pato Nogueira, Marco Vieira · 0 citations
#machine learning Preprint Aug 2026

REPLICANT: Learning Policies for Evading and Hardening Malware Detectors

This work presents Replicant, a deep reinforcement learning framework that learns the realistic task of evasion under a strict label-only black-box threat model and demonstrates that learning the task of evasion not only results in stronger attack performance but provides a better signal for hardening malware detectors...

Shae McFadden, Ilias Tsingenopoulos, Mario D'Onghia et al. · 0 citations

Mitigating Keyword Bias in Java Vulnerability Detection through Dual-Stream CodeBERT with Security Feature Engineering

A dual-stream CodeBERT architecture is presented that addresses keyword bias —by combining pre-trained Transformer representations with a 50-dimensional hand-engineered security feature vector, supported by targeted data augmentation and two-stage adversarial fine-tuning.

Arjun Khurana, Talaya Farasat, Joachim Posegga et al. · 0 citations
#graph neural networks Open access Aug 2026

A Lightweight Hybrid Graph-Neural-Network and Heuristic Framework for Practical Software Vulnerability Assessment in Production Codebases

The significance of this work lies in demonstrating that a deployable, explainable detector can be assembled from compact components, and an edge-type ablation study, a cross-dataset evaluation, and a per-vulnerability analysis are reported to characterize the approach.

A. Elalfy, G. Ebrahim, M. B. Mansour · 0 citations
Preprint Sep 2026

PhantomCall: Evading ML Malware Detectors via Function Call Graph Perturbation

Phan- tomCall is presented, a black-box attack that perturbs the FCG of Windows PE malware by injecting fully executable dummy functions at targeted call sites, adding new nodes and edges to both the CFG and FCG while preserving program semantics.

Md Ajwad Akil, Adrian Shuai Li, Imtiaz Karim et al. · 0 citations
Open access 2026

Malicious Repository Detection Using Transformer-Based Models: A Metadata-Driven Approach

A metadata-driven detection framework that fine-tunes DistilBERT on repository metadata text, specifically the description, README, and topics, and formulates detection as a binary classification task between malware and benign repositories, suggesting that the framework is promising for threshold-based screening scena...

Rabeaa Mouty, M. Abdullah-Al-Wadud · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.