Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents
EvalXRL is a benchmark in which a Large Language Model (LLM) coding agent uses different XRL methods to diagnose a held-out malfunction in an RL agent, and then repair it, and proposes the first head-to-head comparison of multiple XRL methods in closed-loop usage.
Ram Rachum, Yotam Amitai, Bálint Gyevnár et al.
· 0 citations