Aug 2026· Mathematical Modeling and Algorithm Application· Vol 9, pp. 77-83· 0 citations· 10 references
TL;DR
This study uses CL to provide a sample scheduling strategy for RL training, with the core being a three-level adaptive curriculum learning scheduler that dynamically adjusts the mixing ratio of extracting samples of different difficulty levels from the dataset based on the real-time performance of the model during training.
Abstract
Query rewriting is the key to improving the effectiveness of information retrieval, but existing reinforcement learning (RL) based methods often face the challenges of unstable training and slow convergence when the difficulty of the query is uneven. Inspired by the human learning process of "from easy to difficult", curriculum learning (CL), as a general training strategy that can improve the speed of model generalization and convergence, provides new ideas for solving this problem. This study uses CL to provide a sample scheduling strategy for RL training, with the core being a three-level adaptive curriculum learning scheduler. The scheduler dynamically adjusts the mixing ratio of extracting samples of different difficulty levels from the dataset based on the real-time performance (reward) of the model during training, thereby achieving a progressive training strategy for RL models from easy to difficult. Experiments on the HotpotQA dataset show that compared with the random sampling baseline, this method improves the reward value, average performance and training stability by 17.4%, 24.9% and 51.9% respectively, which preliminarily verifies the effectiveness of CL in natural language processing tasks.
The research findings show that RL has gradually expanded from simply improving the accuracy of the final answer to optimizing queries, multi-round search, process decision-making and trustworthy screening, providing new ideas for enhancing the active retrieval ability of RAG and improving the credibility of informatio...
Zun-Long Hong· Applied and Computational En...· 0 citations
DATPO (Difficulty-Adaptive Sentence-entropy-guided Tree-structured Policy Optimization) integrates difficulty-adaptive tree search with a sibling-diversity advantage term, explicitly promoting semantic diversity to expand reasoning coverage during training.
Young Kyu Yu, Sanghwan Jang, Hwanjo Yu· 1 citation
The growing pre-training scale has improved large language models' performance in language generation and knowledge representation. However, their training objectives remain limited to fitting data distributions, thus making it difficult to guarantee that the output satisfies human intentions and safety constraints. Th...
Hong-Yu Jiang· Applied and Computational En...· 0 citations
A new Decentralized Distributed Proximal using Dueling Deep Q Network (D2P-D2QN) is presented, which combines the accuracy of the D2QN estimation with the robustness of proximal policy optimization in a multi-agent setting that is distributed.
Ming Li· Discover Artificial Intellig...· 0 citations
The boundaries of how reinforcement learning optimizes the reasoning capabilities of large language models under different resource and scale constraints are reviewed, as well as the adaptability of reinforcement learning's optimization of large language models in different scenarios are analyzed.
Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}irect Self-\textbf{E}vo...
Yu-Yang Deng, Yu Wang, Jia-Yun Wang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.