Jul 2026· Signal Processing and Communications Applications Conference· pp. 1-4· 0 citations· 12 references
Abstract
General-purpose, low-parameterized language models often produce imprecise or shallow explanations when handling specialized scientific subjects, particularly in non-English contexts like Turkish, where training data is limited. In this paper, an original Turkish biology dataset within the scope of the high school curriculum has been developed; and subsequently, large natural language model studies conducted on this dataset are presented. To synthesize high-quality question-answer pairs, video transcripts and written content were first scraped from public Khan Academy resources. In the two-stage synthesis strategy, by using the GPT-4 API, firstly questions depending on the content and then answers were generated. Data diversity was ensured by producing more than one answer for a single question. The Gemma-3 1B model was fine-tuned on this specialized dataset using the Low-Rank Adaptation (LoRA) method. To measure model performance, LLM-as-a-judge, standard n-gram metrics such as ROUGE, and human evaluation were used. The best-performing configuration, Gemma-3 1B trained on standard-length answers, achieved superior results across all dimensions, reaching M-Prometheus scores of 3.8 for coherence and 3.0 for both correctness and completeness. Additionally, it has been shown that LLM-as-a-judge metrics are closer to human evaluation compared to the ROUGE metric.
Today, Large Language Models (LLMs) perform many tasks in the field of natural language processing with high success, from text generation to translation, semantic analysis to code writing. However, these models have some fundamental limitations that make their reliable use challenging. They can produce factual errors...
M. Toy, Ahmet Ali Süzen· International journal of 3d...· 0 citations
This paper introduces MAAD, a high-quality, carefully constructed and curated by the authors large-scale Arabic dataset for abstractive news summarisation. The authors selected a high-quality subset of 50,000 articles from the dataset Original, which contains 602,792 articles. To maintain the quality, diversity, and tr...
M. Al-Nahari, Ayedh Abdulaziz Mohsen, Nada Abdu Al-Humidi et al.· 2026 6th International Confe...· 0 citations
An in-depth evaluation of instruction tuning for Arabic NLP tasks using three prominent LLMs: LLaMA 3.1-8B, AceGPT-v2-8B, and Qwen3-8B shows that instruction tuning consistently improves performance across most tasks, with notable variations in effectiveness across different tasks and prompts.
Maged Saeed Al-shaibani, Zaid Alyafeai, Irfan Ahmad· Language Resources and Evalu...· 0 citations
MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish, is presented, providing both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.
Uri Katz, Omer Goldman, Tomasz Limisiewicz et al.· 0 citations
This work investigates continued pre-training for adapting large language models to Swedish journalism, using a high-quality dataset that is curate from millions of news articles and demonstrates the importance of targeted evaluation in the adaptation process.
Lukas Borggren, Jenny Kunz, Marco Kuhlmann· 0 citations
This study highlights the role of domain-specific pretraining profile (DSPP) in Transformer performance for modeling digital pragmatics in Arabic-English code-switched discourse. It evaluates MARBERT and XLM-R(oBERTa), with BERT serving as a general-purpose baseline. The models were evaluated on their ability to classi...
Fahad Saud Al Hussen, King Saud University, Riyadh et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.