Skip to content
Review

VulnGym: Benchmarking Coding Agents for Repository-Level Vulnerability Detection

Aug 2026 · 2 citations · 30 references
Computer Science

TL;DR

VulnGym is a real-world repository-level benchmark for evaluating vulnerability detection by coding agents that aligns reviewed GitHub advisories with their corresponding vulnerable version repositories and defines an end-to-end detection task and three oracle-based subtasks to jointly evaluate vulnerability detection and diagnose limitations in code localization and evidence construction.

Abstract

Recent advances in LLM-based vulnerability detection have shown promising results, while coding agents further extend this capability from isolated code snippets to complete repositories. This shift requires agents to autonomously explore repositories and locate vulnerability-relevant code, instead of performing detection on preselected functions. However, existing benchmarks primarily focus on vulnerability classification over preselected code snippets, limiting their ability to evaluate coding agents in repository-level vulnerability detection. Moreover, without fine-grained vulnerability trace annotations, the capability limitations underlying the detection process remain difficult to explore. To address these limitations, we present \textbf{VulnGym}, a real-world repository-level benchmark for evaluating vulnerability detection by coding agents. VulnGym aligns reviewed GitHub advisories with their corresponding vulnerable version repositories. It contains 184 advisories and 408 vulnerability entries across 23 repositories, with each entry annotated with line-level entry points, critical operations, and vulnerability traces. Using this fine-grained ground truth, VulnGym defines an end-to-end detection task and three oracle-based subtasks to jointly evaluate vulnerability detection and diagnose limitations in code localization and evidence construction. Our evaluation indicates that current coding agents remain limited in both end-to-end repository-level vulnerability detection and the construction of accurate supporting traces.

View source

Similar papers

Preprint Aug 2026

VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

VICBench enables robust evaluation of vulnerability detection approaches and shows that state-of-the-art algorithms V-SZZ and LLM4SZZ achieve only 33.3%-40.1% F1, confirming that using existing approaches still entails significant manual effort.

Jin Lu, Xuening Han, Yan Zhong et al. · 0 citations
Open access 2026

From Trace to Line: An Empirical Study of What Drives LLM-Based OSS Vulnerability Localization

This paper introduces T2L (Trace-to-Line), a reproducible research framework that narrows repository-scale code into candidate vulnerable lines through AST-based chunking, structured diagnostic information collection, and evidence-guided refinement that improves trace-to-line localization.

Hao-Ran Xi, Ming-Hao Shao, Brendan Dolan-Gavitt et al. · 0 citations
Jul 2026

VulRESC: A vulnerability detection framework based on risk path extraction and inter-procedural semantic completion

Software vulnerability detection increasingly relies on learning-based models. However, most existing methods analyze individual functions in isolation, making it difficult to capture vulnerabilities caused by cross-function calls; directly introducing complete call chains can also lead to context expansion and noise a...

Yu-Kun Dong, Shuo Wang, Shanchen Pang · 0 citations
#artificial intelligence Preprint Sep 2026

Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale

VLoc Benchmark results establish vulnerability localization as a distinct repository-scale capability and provide a setting for studying both how security agents search for vulnerable code and when they should refrain from reporting it.

Aman Priyanshu, Supriti Vijay, Kimia Majd et al. · 0 citations
Sep 2026

HSF-Vul: hierarchical semantic fusion for vulnerability detection

HSF-Vul is proposed, a novel approach for software vulnerability detection and localization based on hierarchical semantic fusion that frame vulnerability detection as a binary classification task and extend it to line-level localization by analyzing the contribution of individual code lines.

Hong-Tao Wang, Xin Yang, Xiao-Feng Liu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations

LLM-based code generation is now embedded in mission-critical pipelines, but defenses against vulnerable output remain post-hoc -- static analyzers, fine-tuned classifiers, or an LLM judge that screen completed code, ignoring the generating model's own internal state. We test a narrower, directly measurable question: w...

Alizishaan Khatri · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.