Exploring Post-Training Alignment of Small Language Models for Biomedical Data-to-Text Generation: A Case Study of Medication Leaflet
This study presents a comparative analysis of training small language models (SLMs) in specialized biomedical datato-text generation tasks and shows that the aligned SLMs outperform proprietary models like GPT-5; ORPO outperforms the SFTbaselines; and GRPO yields the most robust cross-dataset performance among the alignment methods tested.