This study presents a comparative analysis of training small language models (SLMs) in specialized biomedical datato-text generation tasks and shows that the aligned SLMs outperform proprietary models like GPT-5; ORPO outperforms the SFTbaselines; and GRPO yields the most robust cross-dataset performance among the alignment methods tested.
Abstract
Translating complex biomedical data into patient-friendly narratives is central to modern biomedical informatics. This study presents a comparative analysis of training small language models (SLMs) in specialized biomedical datato-text generation tasks. We explore widely adopted post-training methods including supervised fine-tuning (SFT), direct preference optimization (DPO), odds ratio preference optimization (ORPO), and group relative policy optimization (GRPO) with Qwen-based SLMs on a medicine package leaflets dataset. To assess cross-dataset generalizability, we also curated drug label data from openFDA. We evaluate models using both standard lexical overlap metrics like ROUGE as well as semantic similarity measures. Across our experiments, the results show that (1) the aligned SLMs outperform proprietary models like GPT-5; (2) ORPO outperforms the SFTbaselines; (3) GRPO yields the most robust cross-dataset performance among the alignment methods tested as well as GPT-5.
INTRODUCTION
Knowledge graphs are widely adopted due to their flexible structure and ability to represent large-scale, heterogeneous data. However, as these graphs grow, the queries used to retrieve information become more complex, making efficient information retrieval and question answering more challenging.
OBJECTIVE
This work investigates (1) the impact of few-shot versus zero-shot prompting on Cypher query generation, and (2) performance differences between small (∼7B) and mid-size (∼32B) large language models (LLMs) on both of these approaches.
METHODS
We propose a system that translates natural language questions into Cypher queries for execution on a Neo4j knowledge graph and evaluates answer quality on the PrimeKG and BioHopR datasets. To address data protection and cost constraints, we focus on locally deployable LLMs that require no additional training, using zero-shot and few-shot prompting strategies.
RESULTS
32B models generate syntactically valid Cypher queries using only the knowledge graph schema information. Few-shot prompting with four examples improves F1 score by 35.7 percentage points (∼62% relative improvement) compared to zero-shot prompting. While 7B models struggle to generate syntactically valid queries in the zero-shot setting, few-shot prompting substantially improves their performance, enabling high query syntax validity and F1 scores above 80%.
CONCLUSION
Few-shot prompting heavily improves query generation quality and enables effective use of smaller models, supporting the practical deployment of locally hosted LLMs for knowledge graph question answering.
Suteera Seeha, Adem Abdelmoula, M. Boeker et al.· Studies in Health Technology...· 0 citations
: Biomedical texts present significant challenges for natural language processing (NLP) due to their complex terminology, intricate contextual dependencies, and highly domain-specific semantics. This study investigates the effectiveness of knowledge distillation (KD) for biomedical text classification, aiming to develop lightweight, resource-efficient models that remain competitive with larger architectures. A balanced dataset of 25,000 PubMed records was constructed, equally distributed across five biomedical domains. Two teacher models (BERT and PubMedBERT) and five student models (DistilBERT, BioClinicalBERT, BioBERT, DistilBioBERT, and DistilRoBERTa) were evaluated across ten distinct KD configurations. Each student model was also directly fine-tuned to serve as a controlled baseline. Model performance was assessed using accuracy, precision, recall, and F1-score. The results show that KD can, under suitable teacher–student configurations, enable student models to surpass direct fine-tuning, while other configurations yield only marginal or comparable improvements. Among all configurations, BERT → DistilBERT achieved the highest performance, reaching 89% accuracy. Unexpectedly, the general-purpose BERT teacher outperformed the domain-specific PubMedBERT across multiple student models, suggesting that broader linguistic representations can transfer more effectively across diverse biomedical subdomains. Lower performance in certain KD settings, such as DistilBioBERT and DistilRoBERTa, was attributed to architectural mismatches and limited student capacity. These findings demonstrate that compact models can achieve strong biomedical classification performance through KD under compatible teacher–student pairings, while also highlighting that KD effectiveness varies substantially depending on the specific model combination.
Summary The rapid expansion of biomedical literature demands automated summarization tools that reliably condense research articles into concise, accurate summaries. We benchmarked 62 summarization methods, ranging from frequency-based and TextRank extractors to encoder-decoder models (EDMs) and large language models (LLMs), on 1,000 biomedical abstracts from 20 journals across ScienceDirect and Cell Press, using author-written highlights as reference summaries. Models were evaluated with a composite suite of lexical, semantic, and factual metrics, including ROUGE, BLEU, METEOR, embedding-based similarity, and factuality scores. General-purpose models (e.g., Mistral, GPT, and Llama) achieved the highest overall performance across lexical and semantic dimensions, outperforming reasoning-oriented (e.g., DeepSeek and Magistral) and domain-specific (e.g., BioGPT and BioMistral) models. Notably, medium-sized models outperformed large-scale models, suggesting an optimal balance between model capacity and efficiency, while classical extractive methods lagged behind neural approaches. These findings provide a systematic reference for selecting biomedical summarization tools and highlight that broad pretraining outperforms narrow domain adaptation.
Fabio Baumgärtel, Enrico Bono, Lucas Fillinger et al.· iScience· 0 citations
This study evaluated both closed-source and open-source large language models for extracting and prioritizing medical jargon from EHR notes relevant to individual patients, leveraging prompting techniques, fine-tuning, and data augmentation and found that model performance could deviate largely based on prompting styles.
W. Jang, Sharmin Sultana, Zonghai Yao et al.· JMIR AI· 1 citation
Since the introduction of the Patient Rights Act, patients in Germany have gained legal access to their medical records, including clinical notes. However, these documents are typically written for healthcare professionals and are often difficult for patients to understand due to specialized terminology, abbreviations, and complex sentence structures. Large Language Models (LLMs) offer new opportunities to automatically simplify such texts while preserving medically relevant information. This study investigates the potential of LLMs to improve the readability of German clinical notes by combining a systematic literature review with an experimental evaluation. Ten freely available LLMs were assessed using five synthetic clinical notes, which were simplified through standardized prompts designed to ensure linguistic clarity while maintaining content fidelity. Readability was analyzed using established indices (Flesch Reading Ease, Wiener Sachtextformel, LIX, SMOG, and Coleman-Liau) and complemented by a novel analysis of medical terminology and abbreviation density as indicators of domain-specific complexity. The results show that all LLMs substantially increased overall text length while consistently reducing the density of technical terms and abbreviations. However, no model achieved consistent improvements across all readability indices, highlighting limitations of traditional metrics in the medical domain. Models such as Mistral, ChatGPT, and Copilot demonstrated the highest efficiency in balancing linguistic simplification and text length. Overall, LLMs show strong potential to enhance the accessibility of clinical documentation for patients. However, their effectiveness depends on model selection, prompt design, and evaluation methodology. The findings suggest that conventional readability metrics should be extended with domain-specific measures to more accurately assess comprehensibility in medical texts.
M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al.· Studies in Health Technology...· 0 citations
The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.
Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al.· Cadernos de Saúde Pública· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.