Reinforcement Learning for improving Large Language Models'Catalan text simplification capabilities
A novel reward function is introduced, designed to guide LLMs toward a targeted simplification style with Group Relative Policy Optimization (GRPO), that combines the SARI metric with specific penalty components.
Arnau Ayguadé Domingo, Stefan Bott, Horacio Saggion
· 0 citations