Preprint
Jul 2026
When Synthetic Speech Is All You Have: Better Call GRPO
This work shows that Group Relative Policy Optimization (GRPO) extracts far more from the same synthetic speech than SFT, and traces the gain to behavior rather than representation: GRPO reduces insertion errors by improving stopping calibration and speech-to-text alignment by better anchoring attention to audio, leaving early-layer representations intact.
Shashi Kumar, Yanis Labrak, Hasindri Watawana et al.
· 0 citations