Dataset-Constrained Offline Reinforcement Learning with Value Memory and Preference Feedback
The field of offline reinforcement learning has experienced substantial growth due to its ability to leverage static, previously collected datasets without requiring active exploration in the environment. This paradigm addresses critical safety and sample-efficiency concerns present in online reinforcement learning. Ho...