When2Think: Learning When and How Much to Reason
This work proposes When2Think, an RLVR-based post-training framework for instance-adaptive computation allocation that requires neither a learned reward model nor a learned critic, and offline reference caching avoids online reference-model queries during policy updates.