Safety Hacking in Constrained Best-of-$N$ Inference-time Scaling
It is shown that policies within a bounded $\chi^2$ divergence from the proxy-feasible reference distribution admit an $N$-independent safety-hacking bound, and instantiate this general coverage-control principle with constrained pessimistic sampling.
Akifumi Wachi, Takumi Tanabe, Youhei Akimoto
· 0 citations