Preference benchmarks are built by hiring annotators, and the identity of those annotators is treated as an implementation detail. We measure what that detail buys. On the 2,885 MultiPref items where both pools are internally unanimous, so no tie-breaking convention is consulted at all, expert and crowd annotators assi...
The strongest open-weight coding models are mixture-of-experts (MoE) networks: most of their size comes from large pools of"expert"subnetworks, of which only a few act on any token. That pool is why these models do not fit on the machines most developers own, yet for a user who only wants coding help, most experts enco...
Trusted monitoring has a cheap, trusted model score a stronger untrusted model's actions, and a diverse ensemble of them beats a single stronger monitor at matched cost. They are built by minimising average pairwise correlation, and that paper's twelve monitors shared one base model, leaving open what supplies the dive...
This work separates three arms applied to the same failed candidate: blind whole-solution resampling, spectrum-based localization followed by suspect-span infilling followed by suspect-span infilling, and same-length infilling at a disjoint random code span.
Anik Jha· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.