Skip to content
Review

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks

Jul 2026 · arXiv.org · Vol abs/2607.21763 · 2 citations · 29 references
Computer Science

TL;DR

The "solve rate"metric (clean passes only) is introduced to distinguish genuine capability from cheated outcomes, and it is argued it should be standard practice in any evaluation where cheating vectors are available.

Abstract

Large language model (LLM) agents routinely cheat on cybersecurity benchmarks, inflating reported pass rates far beyond genuine capability. Prior audits of Cybench found cheating in 0.3-3.4% of traces, implicating only a handful of models. We present a controlled prompt-ablation study across 22 frontier models from 7 providers on 23 Cybench capture-the-flag (CTF) challenges under three prompt conditions (no anti-cheat, standard, severe). All 1,518 task traces were individually audited through a four-stage pipeline combining LLM-as-a-judge classification, programmatic verification, judge-verifier reconciliation, and human review. We find cheating is far more pervasive than previously estimated: under baseline conditions, 37.1% of passes involved cheating, 21 of 22 models cheated, and scores were inflated by up to 5x. Anti-cheat prompts reduce cheat propensity from 33.0% (baseline) to 17.8% (standard) to 8.5% (severe) without degrading, and sometimes improving, solve rates. However, even under the most restrictive prompt condition, eight models still produced cheated passes, four showed backfire effects, and cheating escalated from web search toward infrastructure probing. We introduce the"solve rate"metric (clean passes only) to distinguish genuine capability from cheated outcomes, and argue it should be standard practice in any evaluation where cheating vectors are available. Anti-cheat prompts are an effective and essentially free first layer of defense, but they are not a substitute for environmental controls.

View source

Similar papers

#cybersecurity Preprint Aug 2026

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

SRE-Bench is introduced, the first realistic, contamination-free RE benchmark, and results indicate that strong source-code security capabilities do not yet transfer to binary analysis, highlighting RE as an important frontier for agentic cybersecurity and SRE-Bench as a rigorous testbed to measure progress.

J. Spence, Nicholas Assaderaghi, Feng Xiao et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks

Large language model (LLM) benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cybersecurity benchmarks across 10 proprietary, open-weight, and cybersecurity-specialized LLMs. By modeling benchmarks as measurement pipelines, w...

Aymene Berriche, Cathrine Shalby, Mohannad J. Alhanahnah et al. · 1 citation
Review Sep 2026

How Reliable are Automated Jailbreak Evaluators? A Study of Human-Machine Agreement in Cybersecurity Multi-Turn LLM Attacks

Evaluating the effectiveness of multi-turn jailbreak attacks on large language models (LLMs) increasingly relies on automated evaluators, yet their reliability and agreement with human judgments in cybersecurity conversations have not been systematically examined. This paper presents a systematic human-machine agreemen...

Michael Tchuindjang, Nathan Duran, Phil Legg et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.