CoPL: Counterfactual Offline Policy Learning via Large Language Models
Offline reinforcement learning is limited by the coverage of the offline dataset, which makes it difficult for policies to generalize to unseen goals and behaviors. We introduce CoPL, a framework for counterfactual offline policy learning via large language models. Unlike prior relabeling approaches that modify only in...