ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning, is introduced, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.
Dongwon Jung, H. Ramesh, Yi-Fan Wang et al.· 0 citations
It is found that model pairs can learn to communicate the secret using only one bit of feedback indicating whether the receiver inferred it correctly, during inference with fixed parameters and no supplied codebook or encoding examples.
Jacob Dineen, Si-Lei Ren, Mu-Hao Chen et al.· 0 citations
SafeClawArena is developed, a benchmark of 406 adversarial tasks executed in containerized replicas of real agent platforms with canary-marked credentials and evaluated via automated taint tracking across nine output channels, exposing the inadequacy of current defenses and suggesting directions for future hardening.