Frontier coding agents are increasingly trusted to work autonomously for long periods of time, yet what they actually did is often hard to tell from their final response. We quantify the propensity of such agents to overclaim task completion, which may mislead the user. We operationalize overclaiming as a final respons...
This work proposes to use GFlowNet fine-tuning followed by a secondary smoothing phase, to train the attacker model to generate diverse and effective attack prompts, and finds that the attacks generated by the method are effective against a wide range of target LLMs, both with and without safety tuning, and transfer we...
Seanie Lee, Minsu Kim, Lynn Cherif et al.· International Conference on...· 62 citations· ⚡8
A novel unified operator is introduced that combines several regularized RL operators into a general framework that better targets peakier sampling distributions and is named trajectory general mellowmax (TGM), which is shown to identify higher quality, diverse candidates than baselines in both synthetic and real-world...
Marco Jiralerspong, Esther Derman, Danilo Vucetic et al.· 2 citations
This work shows that EFlowNets outperform other GFlowNet formulations in stochastic tasks such as protein design and extends the concept of EflowNets to adversarial environments, proposing adversarial flow networks (A FlowNets) for two-player zero-sum games.
Marco Jiralerspong, Bilun Sun, Danilo Vucetic et al.· International Conference on...· 11 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.