Dataset-Constrained Offline Reinforcement Learning with Value Memory and Preference Feedback
Abstract
The field of offline reinforcement learning has experienced substantial growth due to its ability to leverage static, previously collected datasets without requiring active exploration in the environment. This paradigm addresses critical safety and sample-efficiency concerns present in online reinforcement learning. However, offline reinforcement learning faces severe challenges, predominantly the distributional shift between the behavior policy used for data collection and the learned policy, which often leads to catastrophic extrapolation errors in value estimation. To mitigate these challenges, this paper introduces a novel framework that integrates value memory architectures with preference-based feedback mechanisms within a dataset-constrained optimization landscape. The proposed approach leverages a value memory module to retain high-confidence state-action representations from historical trajectories, anchoring the current value estimates against overestimation biases. Concurrently, human-in-the-loop preference feedback is utilized to regularize the reward landscape, guiding the policy away from spurious correlations inherent in offline datasets. Through rigorous theoretical formulation and empirical validation across continuous control benchmarks, we demonstrate that this dual-mechanism framework significantly improves both policy stability and asymptotic performance. The methodology provides a robust foundation for deploying offline reinforcement learning in complex, real-world scenarios where data is limited and reward functions are difficult to specify analytically.