Sep 2026· Inquiry@Queen's Undergraduate Research Conference Proceedings· 0 citations
TL;DR
Preliminary evidence is found that a single sentence added to the developer's own prompt substantially improves disclosure, and a benchmark scored on attack success alone cannot rank agents on the risk a developer actually carries.
Abstract
Software developers increasingly delegate routine programming work to autonomous AI coding agents that read untrusted project files, run commands, and call external services with limited human review. No developer can audit every action such an agent takes, so oversight defaults to a single artifact: the summary the agent writes when it finishes. The completeness of that summary is a security property in its own right, yet current evaluations do not measure it. Existing benchmarks ask whether a malicious instruction planted in a repository succeeds in redirecting an agent. They do not ask whether the agent then discloses what it did.
This study takes up the second question. We instrument an open-source coding agent running against a controlled, sandboxed repository carrying an injected instruction, and establish per-trial ground truth on what the agent actually did from channels independent of its own account. The design covers nine models, with a second agent used to check that the effect reproduces. Setting that ground truth against the report the developer receives separates two behaviours that prior work has treated as one: whether an attack succeeds, and whether the agent discloses it.
In our data, resisting attack more often does not make a model more forthcoming when an attack does succeed. The two vary independently, so a benchmark scored on attack success alone cannot rank agents on the risk a developer actually carries. We find preliminary evidence that a single sentence added to the developer's own prompt substantially improves disclosure.Faculty Supervisor: Rongxing Lu
This paper evaluated Niyam-AI on 2,000 real-world agent scenarios from Agent-SafetyBench and compared it against three existing safety approaches: NeMo Guardrails, Meta's Llama Prompt Guard 2, and OpenAI's GPT-OSS-Safeguard.
This work presents AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation and introduces AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks.
Wei-Da Liang, Shi Qiu, Zhun Wang et al.· 0 citations
The right unit of safety in agentic AI is the trajectory-with-checkable-evidence, not the model, and this work formalizes an Agent Trajectory Schema and Evidence Chain, state a compositional gating proposition based on standard monitor composition, and outline a research agenda.
Albus W. Ng, Yibin Han, Jusheng Zhang et al.· 1 citation
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skill...
Autonomous coding agents read untrusted files, run shell commands and spawn sub-agents with little supervision, yet their record is usually an editable log. We present Tracekit, an open-source, dependency-free system that captures three channels for every agent session: what the human asked (intent), what the model sai...
A pipeline promoting an AI system publishes records claiming the thing evaluated is the thing deployed and that the evidence licensed the transition, and measures whether those records can express that claim and whether it holds where declared.
Rashid Azarang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.