To Think or Not to Think: Allocating Reasoning Where It Helps
This paper proposes CARE, a method which compares the beneficial length adjustment per question from online sampled responses and applies adaptive length rewards within Group Relative Policy Optimization, with no extra hyperparameters or additional inference cost.
Zheng-Dong He, Yun-Fan Zhou, Jian-Guo Yao et al.
· 0 citations