Hunk-Constrained DPO: Segment-Level Optimization for Secure and Correct LLM Code Generation
Hunk-Constrained Direct Preference Optimization is introduced, a training framework that unifies security hardening and functional correction in large language models and demonstrates that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.