Skip to content
Preprint

LLMVul: A Vulnerability-Labeled Dataset of LLM-Generated C/C++ Functions from Real Production Repositories

Sep 2026 · 0 citations · 23 references
Computer Science

TL;DR

LLMVul, a vulnerability-labeled dataset of LLM-generated C/C++ functions mined from real production repositories, enables reproducible research on vulnerability detection, security evaluation of LLM-generated code, and analysis of vulnerability patterns in AI-assisted software development.

Abstract

Large language models (LLMs) are increasingly used to generate and assist with software development, yet existing vulnerability datasets largely focus on human-written code or controlled prompting environments. This limits the ability to study security weaknesses in LLM-generated code as it appears in real-world software projects. We present LLMVul, a vulnerability-labeled dataset of LLM-generated C/C++ functions mined from real production repositories. We mine AI-assisted development activity from GitHub over a 4 year period, from November 13, 2022 to September 3, 2026, using provenance signals such as commit metadata and AI-related authorship evidence. After filtering and deduplication, LLMVul contains 21,430 unique C/C++ functions from 226 repositories, together with repository, commit, function, provenance, and AI-tool metadata. We establish vulnerability labels using an ensemble of complementary static-analysis and pattern-based techniques and assign Common Weakness Enumeration (CWE) categories to confirmed vulnerable functions. To assess labeling reliability, we additionally conduct independent manual annotation and measure inter-rater agreement using Cohen's kappa ($k=0.79$). LLMVul contains 1,540 ensemble-vulnerable functions spanning 17 unique CWE categories, providing substantially more real-world LLM-generated vulnerable C/C++ functions than existing vulnerability-oriented LLM code benchmarks. By preserving both code-level vulnerability labels and generation/provenance metadata, LLMVul enables reproducible research on vulnerability detection, security evaluation of LLM-generated code, and analysis of vulnerability patterns in AI-assisted software development. The LLMVul dataset is publicly available at https://doi.org/10.5281/zenodo.22668216.

View source

Similar papers

Review

Judging the LLM Judges: A Human-Centric Validation of LLM-Generated Training Data for Software Retrieval

A lightweight quality-assessment protocol is presented for LLM-generated synthetic training data and applied to 13,579 synthetic user reviews generated from GitHub issues across four open-source Android applications, high-lighting the need for hybrid human-AI verification when synthetic data is used in security-critica...

Ogtay Hasanov, Saad Ezzini · 0 citations
Open access 2026

From Trace to Line: An Empirical Study of What Drives LLM-Based OSS Vulnerability Localization

This paper introduces T2L (Trace-to-Line), a reproducible research framework that narrows repository-scale code into candidate vulnerable lines through AST-based chunking, structured diagnostic information collection, and evidence-guided refinement that improves trace-to-line localization.

Hao-Ran Xi, Ming-Hao Shao, Brendan Dolan-Gavitt et al. · 0 citations
Preprint Aug 2026

VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

VICBench enables robust evaluation of vulnerability detection approaches and shows that state-of-the-art algorithms V-SZZ and LLM4SZZ achieve only 33.3%-40.1% F1, confirming that using existing approaches still entails significant manual effort.

Jin Lu, Xuening Han, Yan Zhong et al. · 0 citations
Open access Aug 2026

Multi-SALLM: a multilingual security assessment of generated code

Multi-SALLM, a benchmarking framework designed to systematically evaluate Large Language Models’ ability to generate secure code, reveals three key findings: functional correctness and security are closely related but not equivalent, and sampling strategy is a critical risk factor.

Mohammed Latif Siddiq, Noshin Ulfat, Nishat Raihan et al. · 0 citations
Review Open access 2026

Security Analysis of LLM-Generated Web API Backends

A security assessment on 75 FastAPI backends generated by three contemporary LLMs revealed a disconnect between functional correctness and secure logic, which is interpreted as a review-risk pattern, which is called the human-in-the-loop paradox.

Abdul Ali Khan, S. Rauti, T. Mäkilä · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.