World models (WMs) simulate the transition dynamics of environments, enabling agents to plan over the consequences of their actions. In text-based environments, fine-tuning a Language Model (LM) to serve as a WM has emerged as a dominant paradigm. However, despite the widespread success of non-parametric approaches suc...
Dhananjay Ashok, Shantanu Agarwal, Vivek Datla et al.· 0 citations
Powered by expert guidance, agents can operate in interactive environments; however, it is unclear whether they can learn autonomously from their own experience. To evaluate such self-improvement methods, we introduce GameBoyWorlds, a testbed for agentic self-improvement in video games. GameBoyWorlds-Execution evaluate...
Dhananjay Ashok, Adam Shen, A. Feng et al.· 0 citations
The PAU (Python API Understanding) benchmark is introduced, where models are provided with black-box, API-level access to code snippets, and inspiration is taken from the Asymmetric Actor Critic paradigm, frequently used in robot learning, to post-train models for interactive code understanding.
Dhananjay Ashok, Jesse Thomason, Jonathan May· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.