Skip to content
Open access

Curriculum Learning Query Rewriting Optimization Based on Reinforcement Learning

Aug 2026 · Mathematical Modeling and Algorithm Application · Vol 9, pp. 77-83 · 0 citations · 10 references

TL;DR

This study uses CL to provide a sample scheduling strategy for RL training, with the core being a three-level adaptive curriculum learning scheduler that dynamically adjusts the mixing ratio of extracting samples of different difficulty levels from the dataset based on the real-time performance of the model during training.

Abstract

Query rewriting is the key to improving the effectiveness of information retrieval, but existing reinforcement learning (RL) based methods often face the challenges of unstable training and slow convergence when the difficulty of the query is uneven. Inspired by the human learning process of "from easy to difficult", curriculum learning (CL), as a general training strategy that can improve the speed of model generalization and convergence, provides new ideas for solving this problem. This study uses CL to provide a sample scheduling strategy for RL training, with the core being a three-level adaptive curriculum learning scheduler. The scheduler dynamically adjusts the mixing ratio of extracting samples of different difficulty levels from the dataset based on the real-time performance (reward) of the model during training, thereby achieving a progressive training strategy for RL models from easy to difficult. Experiments on the HotpotQA dataset show that compared with the random sampling baseline, this method improves the reward value, average performance and training stability by 17.4%, 24.9% and 51.9% respectively, which preliminarily verifies the effectiveness of CL in natural language processing tasks.

Read PDF

Similar papers

#large language models Review Open access Sep 2026

Advances in Reinforcement Learning for Retrieval-Augmented Generation in Large Language Model

The research findings show that RL has gradually expanded from simply improving the accuracy of the final answer to optimizing queries, multi-round search, process decision-making and trustworthy screening, providing new ideas for enhancing the active retrieval ability of RAG and improving the credibility of informatio...

Zun-Long Hong · 0 citations
#artificial intelligence Preprint Sep 2026

Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR

DATPO (Difficulty-Adaptive Sentence-entropy-guided Tree-structured Policy Optimization) integrates difficulty-adaptive tree search with a sibling-diversity advantage term, explicitly promoting semantic diversity to expand reasoning coverage during training.

Young Kyu Yu, Sanghwan Jang, Hwanjo Yu · 1 citation
#large language models Open access Sep 2026

Reinforcement Learning Methods and Optimization Strategies in Preference Alignment Techniques for Large Language Models

The growing pre-training scale has improved large language models' performance in language generation and knowledge representation. However, their training objectives remain limited to fitting data distributions, thus making it difficult to guarantee that the output satisfies human intentions and safety constraints. Th...

Hong-Yu Jiang · 0 citations
#reinforcement learning Conference Open access Sep 2026

Progress and Analysis of Optimization for Large Language Models Based on Reinforcement Learning

The boundaries of how reinforcement learning optimizes the reasoning capabilities of large language models under different resource and scale constraints are reviewed, as well as the adaptability of reinforcement learning's optimization of large language models in different scenarios are analyzed.

Feng-Rui Tian · 0 citations
#artificial intelligence Preprint Sep 2026

Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training

Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}irect Self-\textbf{E}vo...

Yu-Yang Deng, Yu Wang, Jia-Yun Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.