Skip to content

BashCoder-R1: Towards Robust and Explainable Bash Script Generation with Robustness-Aware Group Relative Policy Optimization

· 0 citations · 53 references

TL;DR

BashCoder-R1 is presented, a framework that tackles both problems jointly, treating explainability as a design goal rather than a byproduct of correctness, and directly refines the generation policy against a weighted reward combining syntax correctness, robustness as verified by the static analyzer shellcheck, and adherence to the required reasoning format.

View source

Similar papers

Conference Open access 2026

Large Language Model Vulnerabilities

: Large language models are increasingly being deployed in safety-critical domains, yet remain vulnerable to jailbreak attacks that circumvent safety alignments. This systematic review synthesizes empirical jailbreak research published between 2024 and 2025, using a PRISMA-guided search protocol, followed by BERTopic-b...

Meda Račaitytė, Hélder Bastos, R. Ribeiro et al. · 0 citations
#natural language process... Preprint Sep 2026

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving functional correctness, these agents must faithfully follow process instructions and c...

Bo-Si Wen, Cun-Xiang Wang, Jia-Yi Gui et al. · 0 citations
Open access Aug 2026

CodeVulReason: Incentivizing reasoning for code vulnerability detection

This paper introduces CodeVulReason, a unified framework for enhancing the reasoning capabilities of large language models (LLMs) in code vulnerability detection (CVD), and proposes S-LoRA, a parameter-efficient fine-tuning method that optimizes Low-Rank Adaptation (LoRA) rank allocation through a statistically grounde...

Zhengye Li, Kenny Zhu · 0 citations

Hunk-Constrained DPO: Segment-Level Optimization for Secure and Correct LLM Code Generation

Hunk-Constrained Direct Preference Optimization is introduced, a training framework that unifies security hardening and functional correction in large language models and demonstrates that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.

Qian-Shuo Huang, Xin Yin, Xin-Rui Li et al. · 0 citations
Review Aug 2026

Refine After Generation: Toward Correct and Concise Patches in LLM-based Program Repair

This paper identifies patch verbosity as a major yet overlooked concern in LLM-based APR and proposes RECAP, a lightweight, plug-and-play adapter that attaches to existing repair frameworks after generation that achieves a substantially better size-correctness tradeoff.

Wen-Qiang Luo, J. Keung, Xiaoyu Shi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.