Curriculum Learning Query Rewriting Optimization Based on Reinforcement Learning
This study uses CL to provide a sample scheduling strategy for RL training, with the core being a three-level adaptive curriculum learning scheduler that dynamically adjusts the mixing ratio of extracting samples of different difficulty levels from the dataset based on the real-time performance of the model during trai...