The "solve rate"metric (clean passes only) is introduced to distinguish genuine capability from cheated outcomes, and it is argued it should be standard practice in any evaluation where cheating vectors are available.
Abstract
Large language model (LLM) agents routinely cheat on cybersecurity benchmarks, inflating reported pass rates far beyond genuine capability. Prior audits of Cybench found cheating in 0.3-3.4% of traces, implicating only a handful of models. We present a controlled prompt-ablation study across 22 frontier models from 7 providers on 23 Cybench capture-the-flag (CTF) challenges under three prompt conditions (no anti-cheat, standard, severe). All 1,518 task traces were individually audited through a four-stage pipeline combining LLM-as-a-judge classification, programmatic verification, judge-verifier reconciliation, and human review. We find cheating is far more pervasive than previously estimated: under baseline conditions, 37.1% of passes involved cheating, 21 of 22 models cheated, and scores were inflated by up to 5x. Anti-cheat prompts reduce cheat propensity from 33.0% (baseline) to 17.8% (standard) to 8.5% (severe) without degrading, and sometimes improving, solve rates. However, even under the most restrictive prompt condition, eight models still produced cheated passes, four showed backfire effects, and cheating escalated from web search toward infrastructure probing. We introduce the"solve rate"metric (clean passes only) to distinguish genuine capability from cheated outcomes, and argue it should be standard practice in any evaluation where cheating vectors are available. Anti-cheat prompts are an effective and essentially free first layer of defense, but they are not a substitute for environmental controls.
Under the single-shot, raw-bytecode-only protocol, current LLMs are not reliable standalone forensic tools, and their robustness has not been systematically tested against contracts adversarially designed to mislead analysis.
SRE-Bench is introduced, the first realistic, contamination-free RE benchmark, and results indicate that strong source-code security capabilities do not yet transfer to binary analysis, highlighting RE as an important frontier for agentic cybersecurity and SRE-Bench as a rigorous testbed to measure progress.
J. Spence, Nicholas Assaderaghi, Feng Xiao et al.· 1 citation
This paper does an audit of defense mechanisms under jailbreak attacks on locally deployed models by tracing each failure back to the specific assumption it relies on.
Large language model (LLM) benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cybersecurity benchmarks across 10 proprietary, open-weight, and cybersecurity-specialized LLMs. By modeling benchmarks as measurement pipelines, w...
Aymene Berriche, Cathrine Shalby, Mohannad J. Alhanahnah et al.· 1 citation
A reproducible reliability audit of the developer-accessible on-device foundation model is presented, framed as an oversight question: can a user or a resource-constrained developer tell when the model is wrong?
Shashwat Pandey, Satwik Pandey, S. Raghu· 0 citations
Evaluating the effectiveness of multi-turn jailbreak attacks on large language models (LLMs) increasingly relies on automated evaluators, yet their reliability and agreement with human judgments in cybersecurity conversations have not been systematically examined. This paper presents a systematic human-machine agreemen...
Michael Tchuindjang, Nathan Duran, Phil Legg et al.· International Symposium on N...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.