Skip to content
Conference Open access

Disclosure Failure in Compromised AI Coding Agents

Sep 2026 · Inquiry@Queen's Undergraduate Research Conference Proceedings · 0 citations

TL;DR

Preliminary evidence is found that a single sentence added to the developer's own prompt substantially improves disclosure, and a benchmark scored on attack success alone cannot rank agents on the risk a developer actually carries.

Abstract

Software developers increasingly delegate routine programming work to autonomous AI coding agents that read untrusted project files, run commands, and call external services with limited human review. No developer can audit every action such an agent takes, so oversight defaults to a single artifact: the summary the agent writes when it finishes. The completeness of that summary is a security property in its own right, yet current evaluations do not measure it. Existing benchmarks ask whether a malicious instruction planted in a repository succeeds in redirecting an agent. They do not ask whether the agent then discloses what it did. This study takes up the second question. We instrument an open-source coding agent running against a controlled, sandboxed repository carrying an injected instruction, and establish per-trial ground truth on what the agent actually did from channels independent of its own account. The design covers nine models, with a second agent used to check that the effect reproduces. Setting that ground truth against the report the developer receives separates two behaviours that prior work has treated as one: whether an attack succeeds, and whether the agent discloses it. In our data, resisting attack more often does not make a model more forthcoming when an attack does succeed. The two vary independently, so a benchmark scored on attack success alone cannot rank agents on the risk a developer actually carries. We find preliminary evidence that a single sentence added to the developer's own prompt substantially improves disclosure.Faculty Supervisor: Rongxing Lu

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

This work presents AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation and introduces AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks.

Wei-Da Liang, Shi Qiu, Zhun Wang et al. · 0 citations
Review Aug 2026

Agent Safety Should Be a Runtime Contract

The right unit of safety in agentic AI is the trajectory-with-checkable-evidence, not the model, and this work formalizes an Agent Trajectory Schema and Evidence Chain, state a compositional gating proposition based on standard monitor composition, and outline a research agenda.

Albus W. Ng, Yibin Han, Jusheng Zhang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skill...

T. Kaiser, Aritra Dhar · 0 citations
Review Sep 2026

Tracekit: Tamper-Evident Intent-Reasoning-Action Auditing for Autonomous Coding Agents

Autonomous coding agents read untrusted files, run shell commands and spawn sub-agents with little supervision, yet their record is usually an editable log. We present Tracekit, an open-source, dependency-free system that captures three channels for every agent session: what the human asked (intent), what the model sai...

Bravish Ghosh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.