Preprint
Jul 2026
ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm
Evaluating frontier CLI agents, it is found that while they often refuse illegal tasks when prompted directly, compliance reaches 100\% under persistent malicious interaction, and it is demonstrated that current alignment techniques are insufficient for autonomous agents.
Kefan Song, Yan-Jun Qi
· 0 citations