Item Response Theory for AI Safety
Overall, it is shown IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which frontier labs and evaluators adopt.
We have 2 of 9 papers
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
Overall, it is shown IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which frontier labs and evaluators adopt.
The Activation Controllability Benchmark is introduced to quantify the extent to which models can modulate their residual stream via natural-language instruction, and suggests that control over the activation space itself could become a confound for monitoring as introspective capabilities increase.
We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.