Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
It is found that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.