Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks

As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Measuring reward hacking is itself challenging, as detection typically relies on human inspection or LLM judges, both of which can be unreliab...

Amit Roth, I. Bercovich, Yonathan Efroni · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.