Skip to content

Hunk-Constrained DPO: Segment-Level Optimization for Secure and Correct LLM Code Generation

Sep 2026 · ACM Transactions on Software Engineering and Methodology · 0 citations · 44 references
Advanced Malware Detection Techniques Software Engineering Research

TL;DR

Hunk-Constrained Direct Preference Optimization is introduced, a training framework that unifies security hardening and functional correction in large language models and demonstrates that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.

Abstract

Large Language Models (LLMs) have been widely applied in code generation tasks like code completion and automated development, demonstrating significant potential for improving coding efficiency. However, research has shown that LLM-generated code frequently contains security vulnerabilities, raising concerns about its reliability in production environments. To address these security issues, various mitigation approaches have been proposed, but these methods typically impact the LLM’s ability to generate functionally correct code, which may limit their practical application in real-world development environments. In this work, we address this problem through a key observation: security patches and functional bug fixes in real-world software exhibit structural similarities as small, localized modifications. This shared characteristic suggests that a unified learning model could address both objectives jointly. Building on this insight, we introduce HPO (Hunk-Constrained Direct Preference Optimization), a training framework that unifies security hardening and functional correction. Our framework features two key technical components: a novel segment-weighted preference optimization objective to focus learning on repair logic, and an automated data synthesis pipeline to provide high-quality training data. Experiments across multiple models and programming languages demonstrate that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.

View source

Similar papers

Preprint Open access Aug 2026

Understanding and Improving Model Editing for Secure Code Generation

The first systematic study of model editing as a model-level hardening mechanism for secure code generation is conducted, evaluating 3 state-of-the-art editing methods across diverse LLM families and comparing them with CoSec, a representative inference-time approach, focusing on security, robustness, generalization, a...

Wei-Feng Sun, Quan-Jun Zhang, Yuchen Chen et al. · 0 citations
#machine learning Preprint Sep 2026

Introspective Uncertainty Estimation for LLM-Based Code Generation

The findings suggest that hidden states are a robust and informative resource for estimating functional code correctness, supporting a two-stage workflow that combines response-level risk screening with targeted line-level prioritization.

T. Klassert · 0 citations
Open access Aug 2026

Multi-SALLM: a multilingual security assessment of generated code

Multi-SALLM, a benchmarking framework designed to systematically evaluate Large Language Models’ ability to generate secure code, reveals three key findings: functional correctness and security are closely related but not equivalent, and sampling strategy is a critical risk factor.

Mohammed Latif Siddiq, Noshin Ulfat, Nishat Raihan et al. · 0 citations
#small language model Preprint Sep 2026

Code Transformation Rule Synthesis using LLMs: Potential and Limits

Due to their black-box nature, LLMs suffer from limited explain- ability and a lack of determinism. Their usage cost can also rise, particularly with repetitive tasks on large codebases. To mitigate this, we conduct a novel empirical study targeting three domain- specific languages for transformation rules, namely Comb...

Axel Allain, Aymeric Blot, D. Khelladi et al. · 1 citation
#artificial intelligence Preprint Sep 2026

JevVibe: Efficient Classification-Guided Secure Code Generation

This work evaluates Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, and builds JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code...

Arshak Rezvani, Sasha Behrouzi, Ahmad Sadeghi · 0 citations
Preprint Aug 2026

Route-Align-Verify for Functional Correctness in Code Generation

The results indicate that functional correctness in code generation can be meaningfully improved without modifying the backbone architecture, by jointly optimizing how tasks are prompted, how the model is adapted, and how final outputs are selected.

Eric Zhou, Jing Meng, Ao-Fan Liu · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.