Sprint or Delve: A Distribution-Aware Approach to Efficient Reasoning
The Powered Length Penalty (PLP) is proposed, an adaptive regularizer that penalizes redundancy in short sequences while gradually reducing penalties for longer sequences, preserving deep reasoning.