Conference
Open access
2026
The Inner Monologue of Language Models: When Reasoning Traces Reveal More Than They Hide
It is indicated that RL-trained models not only demonstrate greater awareness of their learned behaviors and stronger generalizability to novel, structurally similar tasks than SFT models but also often exhibit weak alignment between their reasoning traces and final outputs, an effect most pronounced in GRPO-trained models.
Pratham Singla, Shivank Garg, Ayush Singh et al.
· Annual Meeting of the Associ... · 0 citations