This work evaluates Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, and builds JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code generated by Qwen2.5-Coder-32B-Instruct.
Abstract
Large language models can generate functionally correct code that still contains security weaknesses, motivating repair pipelines that first diagnose a weakness type before deciding how to fix it. The Common Weakness Enumeration (CWE) provides a standardized vocabulary for such diagnoses, but asking an autoregressive language model to generate a CWE label and extracting it from the response raises questions about output validity, speed, and cost, as well as accuracy. We evaluate Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, GPT-5.6-Sol, on a controlled 50-way CWE classification task over 1,916 CyberSecEval benchmark examples. Jev outperforms all six open-weight baselines on every classification and ranking metric, while its comparison with GPT-5.6-Sol depends on the metric: GPT-5.6-Sol achieves higher Top-1 accuracy and Macro-F1, whereas Jev achieves higher Top-3 and Top-5 accuracy and a nearly identical MRR, at $6.27\times$ lower median API latency and $55.9\times$ lower estimated API cost. We further build JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code generated by Qwen2.5-Coder-32B-Instruct. With Jev providing the diagnosis, the agent increases the detector-measured security pass rate from 63.5% before repair to 70.7%, compared with 66.1% for LLM-guided repair. These results show that JevVibe is effective at improving the security of generated code, with Jev providing reliable and efficient CWE classification.
Hunk-Constrained Direct Preference Optimization is introduced, a training framework that unifies security hardening and functional correction in large language models and demonstrates that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.
Qian-Shuo Huang, Xin Yin, Xin-Rui Li et al.· ACM Transactions on Software...· 0 citations
LLM-based code generation is now embedded in mission-critical pipelines, but defenses against vulnerable output remain post-hoc -- static analyzers, fine-tuned classifiers, or an LLM judge that screen completed code, ignoring the generating model's own internal state. We test a narrower, directly measurable question: w...
Large language models (LLMs) are increasingly used for code generation, yet they remain vulnerable to prompts that elicit insecure implementations. Existing defenses typically rely on predefined threat models or known vulnerability patterns, limiting their effectiveness against novel attacks. We propose CodeSIFT, a thr...
Francesco Quinzan, Noor Munir, Yi-Shun Lu et al.· 0 citations
These findings show that reliable evaluation of LLM-generated code requires validated ground truth, protected tests, and multiple explicitly interpreted measures, and that CodeAssay provides a reproducible basis for evidence-based model evaluation in AI-augmented software development.
Shahbaz Siddeeq, Muhammad Waseem, Umar Subhan Malhi et al.· 0 citations
The first systematic study of model editing as a model-level hardening mechanism for secure code generation is conducted, evaluating 3 state-of-the-art editing methods across diverse LLM families and comparing them with CoSec, a representative inference-time approach, focusing on security, robustness, generalization, a...
Wei-Feng Sun, Quan-Jun Zhang, Yuchen Chen et al.· 0 citations
Multi-SALLM, a benchmarking framework designed to systematically evaluate Large Language Models’ ability to generate secure code, reveals three key findings: functional correctness and security are closely related but not equivalent, and sampling strategy is a critical risk factor.
Mohammed Latif Siddiq, Noshin Ulfat, Nishat Raihan et al.· International Conference on...· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.