Conference
Open access
Sep 2026
Disclosure Failure in Compromised AI Coding Agents
Preliminary evidence is found that a single sentence added to the developer's own prompt substantially improves disclosure, and a benchmark scored on attack success alone cannot rank agents on the risk a developer actually carries.
Aaron Su, Rong-Xing Lu
· Inquiry@Queen's Undergraduat... · 0 citations