Skip to content

Open-Source Intelligence for Code Provenance and the Security Patterns that Separate Human and Large-Language-Model Implementations of Common Programming Tasks

Jul 2026 · arXiv.org · Vol abs/2607.12524 · 0 citations · 16 references
Computer Science

TL;DR

A fully reproducible pipeline is built that collects real implementations of 31 security-sensitive programming tasks, among them OAuth with PKCE, JWT verification, password hashing, and SQL access, from 9 language models and from human answers, and scores every sample with deterministic security and style detectors.

Abstract

Developers now draw code from two very different sources, the accumulated human answers on sites such as Stack Overflow and the output of large language models. We ask two questions about that split. First, can the provenance of a code snippet be recovered from the code itself, and second, do the two sources differ in the security patterns they adopt for the same task. Using only open sources, a public gateway of open-weight language models and the public Stack Overflow API, we build a fully reproducible pipeline that collects real implementations of 31 security-sensitive programming tasks, among them OAuth with PKCE, JWT verification, password hashing, and SQL access, from 9 language models and from human answers, and scores every sample with deterministic security and style detectors. On 528 real samples we train a cross-validated classifier that recovers human versus model provenance with 93 percent accuracy against a 78 percent baseline, and a 7-way classifier that attributes a sample to the specific model that wrote it at 48 percent. We then report where the sources diverge on security, which patterns models adopt more often than the human corpus and which they inherit from it. Running the same tasks in Python, JavaScript, Go, and Java, we find the security divergence holds in every language while the provenance boundary is partly language-specific and does not transfer symmetrically between them. A vulnerability repair case study, in which models are handed insecure code and asked to fix it, finds a 77 percent repair rate across 21 seeds and 12 weakness classes, but a recurring partial-fix failure in which the model removes the insecure pattern without adding the correct defense. The pipeline is data driven, so any new task or language is added as a single specification entry, and a fail-closed checker re-derives every number in this paper from the stored data.

View source

Similar papers

Preprint Sep 2026

LLMVul: A Vulnerability-Labeled Dataset of LLM-Generated C/C++ Functions from Real Production Repositories

LLMVul, a vulnerability-labeled dataset of LLM-generated C/C++ functions mined from real production repositories, enables reproducible research on vulnerability detection, security evaluation of LLM-generated code, and analysis of vulnerability patterns in AI-assisted software development.

Mohammad Farhad, Shuvalaxmi Dass · 0 citations
2026

A Code-Refactoring-Based Migration Scheme for Trusted Applications across Heterogeneous TEEs

Experimental results show that CRTAMS enables trusted-application migration across heterogeneous TEEs with low refactoring cost and moderate execution overhead in the evaluated workloads, while reducing the amount of platform-specific code that developers must write manually.

Di Lu, Qing-Wen Zhang, Yujia Liu et al. · 0 citations
#small language model Preprint Aug 2026

Vulnerable Code Search: Transferable Attack for Code Language Models

This paper introduces a programming language-agnostic, transferable, adversarial attack that exploits this CLM vulnerability and demonstrates that this attack, even when computed using smaller code embedding models, is highly effective and transferable to larger, closed-source embedding models.

Kaicheng Wang, Liyan Huang, Jesse Thomason et al. · 0 citations
Open access Sep 2026

On-Premise CodeBERT-Driven Model for Vulnerability Detection in Source Code

An on premise Artificial Intelligence (AI)-based model for vulnerability detection in source code, designed to ensure there is efficiency in identifying potential weaknesses, and deployed locally within a Dockerized environment.

D. Sako · 0 citations
Book Open access Aug 2026

Towards Lightweight Domain Model Reverse Engineering from Source Code

This paper proposes an automated approach to extract domain models from source code using lightweight, locally deployable LLMs and achieves high F1-scores on a dataset of ten projects, each comprising a curated domain model and its corresponding implementation, while remaining fully executable on locally deployable LLM...

Kévin Delcourt, Meriem Ben Chaaben, Abdelhamid Rouatbi et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.