Skip to content
Conference

Beyond Lexical Overlap: A Multi-Dimensional Evaluation Framework for Arabic Retrieval-Augmented Generation Systems

Aug 2026 · 2026 6th International Conference on Emerging Smart Technologies and Applications (eSmarTA) · pp. 1-6 · 0 citations · 32 references

Abstract

Retrieval-Augmented Generation (RAG) has become a standard technique to ground large language model outputs in external knowledge. However, evaluating RAG systems for Arabic remains problematic because traditional lexical metrics such as ROUGE and BLEU were designed for English, a language with limited morphological variation. Arabic's rich derivational system where a single triliteral root can produce dozens of surface forms causes these metrics to penalize correct paraphrasing while being completely bling to factual hallucination. This paper introduces a holistic evaluation framework that combines hybrid retrieval (BM25 with dense embeddings from BGE-M3) and a multi-dimensional scoring suite. We compare six LLMs three Arabic specialized (ALLaM-7B, Fanar-1-9B, Noon-7b) and three multilingual (Llama-3.3-70B, Qwen-2.5-7B, BLOOM-7B) across 300 queries over a general Arabic corpus of 30 documents. Our hybrid retriever achieves aRecall@5 of 0.942 and an MRR of0.883. The central empirical finding is a metric mirage: the Pearson correlation between ROUGE-1 and semantic similarity (measured by BGE-M3) is only 0.317, meaning that more than 90% of factual variance is invisible to lexical overlap metrics. Llama-3.3-70B achieves the highest composite score (2.703/4), driven by superior named entity recognition accuracy (0.463). However, the smaller Arabic-specialized ALLaM-7B exhibits strong performance across several of evaluation metrics—particularly named entity recognition (NER), ROUGE-1, ROUGE-2, and BLEU. https://github.com/AAA20121/Beyond-Lexical-Overlap.git.

View source

Similar papers

2026

Efficient Adaptation of English Language Models for Morphologically Rich and Underrepresented Languages: The Case of Arabic

A resource-efficient adaptation of the English-pretrained ModernBERT for Arabic, employing continued pretraining on large Arabic corpora followed by lightweight head-only fine-tuning with a frozen encoder, demonstrating that modern English encoder architectures can be efficiently transferred to Arabic through language-...

Ahmed Samy Eldamaty, M. Abdelrahman, Mohamed Mostafa Ibrahim Elbehery et al. · 0 citations
Preprint Aug 2026

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study

This paper proposes a general optimization framework that combines a vocabulary pruning method with a targeted fine-tuning protocol for MNMT models, and reduces the vocabulary size from over 128,000 to approximately 10,000 tokens, enabling a 60% memory saving without any loss in performance.

A. A. Aliane, N. Semmar, H. Aliane · 0 citations
#natural language process... Preprint Sep 2026

AraGenre 2026: A Hierarchical Definition-Guided Arabic Genre Classification Shared Task

AraGenre is a shared task on hierarchical, definition-guided Arabic genre classification, motivated by the limited availability of annotated data in Arabic and other low-resource languages. Systems assign each Arabic text segment both a broad communicative genre and a fine-grained specific genre. The released training...

Mo El-Haj, Saad Ezzini, Shadi Abudalfa et al. · 0 citations
Aug 2026

STAR: instruction tuning for Arabic across tasks, datasets, and models

An in-depth evaluation of instruction tuning for Arabic NLP tasks using three prominent LLMs: LLaMA 3.1-8B, AceGPT-v2-8B, and Qwen3-8B shows that instruction tuning consistently improves performance across most tasks, with notable variations in effectiveness across different tasks and prompts.

Maged Saeed Al-shaibani, Zaid Alyafeai, Irfan Ahmad · 0 citations
Open access Sep 2026

Design and Implementation of a Scalable AI-Based Semantic Evaluation System for Hindi Text Using Transformer Models

Evaluating linguistically diverse descriptive answers in a consistent and accurate manner in modern digital education systems is a growing challenge, especially in low-resource languages like Hindi. Traditional lexical and rule-based grading systems cannot adequately reflect the meaning behind the words, negation, para...

Nirja D. Shah, Jyoti Pareek · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.