FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations
This work introduces Few-Shot Aware GRPO (FSA-GRPO), an RL-based post-training recipe that uses a specially designed reward to encourage the model to leverage few-shot demonstrations, thereby strengthening its few-shot adaptation ability.