Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
We study training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different probabilities to the same tokens. To account for this discrep...