Skip to content

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning

Jul 2026 · arXiv.org · Vol abs/2607.27610 · 2 citations · ⚡ 2 influential · 38 references
Computer Science

TL;DR

A Kalman-Guided Prompt Selection method (KGPS), which reformulates prompt selection as a dynamic state estimation problem rather than static difficulty prediction, and consistently improves both final accuracy and rollout efficiency over strong baselines, establishing state-of-the-art performance among online prompt selection methods.

Abstract

Reinforcement learning (RL) finetuning significantly enhances the reasoning capabilities of large language models (LLMs), yet its effectiveness critically depends on selecting prompts of appropriate difficulty for the current policy. This is challenging because prompt difficulty evolves throughout training. Existing online methods therefore face a trade-off: evaluation-based approaches are accurate but expensive, while prediction-based approaches are efficient but typically assume stationary difficulty, making them ill-suited to RL's non-stationary training dynamics. To address these issues, we propose a Kalman-Guided Prompt Selection method (KGPS), which reformulates prompt selection as a dynamic state estimation problem rather than static difficulty prediction. KGPS models each prompt's latent success rate in logit space using a linear-Gaussian state-space model, with process noise coupled to the magnitude of policy updates so that uncertainty increases when the policy changes more substantially. A Kalman filter then maintains a calibrated Gaussian posterior over prompt difficulty, and prompts are selected by maximizing a posterior-expected training utility that favors intermediate-difficulty prompts while naturally revisiting uncertain ones. The resulting procedure is adaptive to policy drift and requires no additional rollouts beyond standard policy training. Extensive experiments across mathematics, planning, and geometry reasoning benchmarks, as well as multiple RL algorithms, show that KGPS consistently improves both final accuracy and rollout efficiency over strong baselines, establishing state-of-the-art performance among online prompt selection methods. For example, on DeepSeek-R1-Distill-7B, KGPS uses 83% fewer rollouts than DS while even improving the average performance by 0.12 point across six math reasoning benchmarks.

View source

Similar papers

#machine learning Preprint Aug 2026

PAC: Progress-Augmented Advantage Curriculum for Multi-Task Reinforcement Learning of LLMs

PAC, a Progress-Augmented Advantage Curriculum for multi-task RL of LLMs that combines two task-level signals: advantage-derived learnability, which measures the magnitude of the policy update a task can induce, and recent reward gains, which show whether those updates have improved task performance.

Yuan-Qiang Yu, Yan-Zhao Zheng, Zhen-Tao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration

Reinforcement learning with verifiable rewards (RLVR) is often limited by insufficient exploration: difficult problems can yield uniformly incorrect rollout groups and therefore little learning signal. We show that such failures need not reflect missing capability. Instead, finite sampling often concentrates on a probl...

Jin Cui, Xin-Yue Long, Bo-Ran Zhao et al. · 0 citations
#natural language process... Preprint Sep 2026

Expert-Space Exploration in MoE Reinforcement Learning

Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the s...

Hong-Yi He, Zheng-Wen Lin, Xiao Liu et al. · 0 citations
#machine learning Preprint Sep 2026

MaPP: A Unified Marginalized Posterior-Predictive Framework for Data-Efficient RLVR

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but incurs substantial costs from rollouts and policy updates. Online prompt selection improves efficiency by using per-prompt Bayesian posteriors to predict difficulty and prioritize informative prompts....

Yang-Yang Ren, Hao-Dong Zhu, Sheng Xu et al. · 0 citations
#machine learning Preprint Sep 2026

Score Centering Stabilizes Off-policy Reinforcement Learning

It is shown that the instability of RL under TIM is primarily caused by drift: a persistent bias between training and inference engines that accumulates with every training step, and an additive score centering term is derived that stabilizes RL under TIM by canceling drift.

Martina Marek, Max Ryabinin · 2 citations · ⚡1
Preprint Aug 2026

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning

This work proposes CVPO - Curriculum-guided Value-Variance Policy Optimization, a dynamic curriculum weighting method that adapts to question difficulty that achieves better performance and stronger exploration, enabling more accurate and robust reasoning in language models across various math tasks.

Ziqi Jia, Yalu Ouyang, Bo Pang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.