Skip to content
Preprint

Evaluating and Preventing Security Smells in AI-Generated Ansible Code

Aug 2026 · 0 citations · 28 references
Computer Science

TL;DR

An approach integrating Ansible best practices and CIS benchmarks into prompts through an extended CO-STAR framework is introduced, enabling security smell prevention during synthesis rather than detection after deployment, leading to a fourfold improvement over humans at 23%-43%, with overall code quality improving by 19%-49%.

Abstract

AI coding assistants generate Infrastructure as Code, yet no work has examined whether this code meets security requirements. This matters because security smells in infrastructure code propagate to deployed systems, producing infrastructure that is insecure and untrustworthy. We evaluate 16 AI models generating Ansible roles for Apache Tomcat v10 and MongoDB v7, analysing 278 Ansible roles against CIS benchmarks. Without security guidance, all 16 AI models produced code containing security smells, resulting in vulnerable infrastructure that fails compliance verification and underperforms code written by human developers. We introduce an approach integrating Ansible best practices and CIS benchmarks into prompts through an extended CO-STAR framework, enabling security smell prevention during synthesis rather than detection after deployment. When this approach is applied, 4 out of 16 models generate compliant code, with the leading model achieving 95%-100% CIS compliance, a fourfold improvement over humans at 23%-43%, with overall code quality improving by 19%-49%. The remaining 12 models fail not because they cannot generate code but because they cannot follow instructions with multiple constraints. For capable models, the approach requires no retraining and can be adopted through system prompts.

View source

Similar papers

Preprint Aug 2026

Securing AI-Generated Code: A Just-in-Time Vulnerability Detection and Remediation Pipeline

An automated security evaluation pipeline that generates Python code from LLMSecEval prompts, scans for vulnerabilities using CodeQL and Bandit in parallel with an independent Code Validator LLM, enriches the Code Validator findings with MITRE ATT&CK techniques, CWE Observed Examples, and Python best practice guideline...

Mikhail Surikov · 0 citations
Open access 2026

Shift-Left Security for AI-Generated Code: Detecting and Preventing Vulnerabilities at Build-Time

A structured, research-backed blueprint for build-time security controls tailored specifically to AI-generated code is developed, organized across three interdependent layers: a pre-pull-request policy and governance layer; a build-time detection layer combining AI-augmented and traditional static analysis; and a prior...

Bala Thripura Akasam · 0 citations
#artificial intelligence Preprint Aug 2026

Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code

GenIaC-SecBench is introduced, a benchmark of 100 deployment scenarios stratified by architectural complexity, evaluated across 12 model configurations from four vendors, producing 1,196 IaC artifacts scanned by three independent policy engines (Checkov, Trivy, KICS).

Animesh Shaw · 0 citations
Jul 2026

CHARGE: Leveraging CWE Hierarchies for Hardware Security SystemVerilog Assertion Generation

CHARGE is an automated framework for generating security properties for unverified RTL modules using CWEs and large language models using CWEs and large language models that leverages the hierarchical nature of CWE entries to improve accuracy when identifying security-critical assets in unverified RTL modules.

Xiao Tan, C. Sturton · 0 citations
Review

Security Vulnerabilities in AI-Generated JWT Authentication Code for Spring Boot

Investigating the security vulnerabilities present in AI-generated JWT authentication code for Java Spring Boot Representational State Transfer Application Programming Interfaces (REST API) reinforces that AI-generated JSON Web Token (JWT) authentication code requires dedicated security review.

Hoang Long Nguyen, Mezid Hmudda, Benjamin Powley · 0 citations
Conference Jul 2026

Is AI-Generated Web Code Vulnerability-Free?

With the increasing usage of AI-generated code in software development workflows, new security challenges and concerns arise. This paper analyzes five LLMs: ChatGPT, Claude, Gemini, DeepSeek, and Grok in three phases of security assessments against web vulnerabilities listed by the OWASP Top 10. Phase 1 (December 2025)...

Malak Mansour, Anas AlMajali · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.