Skip to content

Benchmarking and AI-assisted human-like evaluation of retrieval-augmented generation for Arabic and English documents

Jul 2026 · Language Resources and Evaluation · Vol 60 · 0 citations · 47 references
Computer Science

TL;DR

An important practical reproducible framework for multilingual RAG benchmarking and insights for optimizing performance on resource-constrained devices are contributed and support the latent language hypothesis by suggesting an internal model bias toward high-resource languages.

View source

Similar papers

Conference Aug 2026

Beyond Lexical Overlap: A Multi-Dimensional Evaluation Framework for Arabic Retrieval-Augmented Generation Systems

Retrieval-Augmented Generation (RAG) has become a standard technique to ground large language model outputs in external knowledge. However, evaluating RAG systems for Arabic remains problematic because traditional lexical metrics such as ROUGE and BLEU were designed for English, a language with limited morphological va...

Ahmed Ali Al-Ansi, Khalil Al-Wagih · 0 citations
Aug 2026

STAR: instruction tuning for Arabic across tasks, datasets, and models

An in-depth evaluation of instruction tuning for Arabic NLP tasks using three prominent LLMs: LLaMA 3.1-8B, AceGPT-v2-8B, and Qwen3-8B shows that instruction tuning consistently improves performance across most tasks, with notable variations in effectiveness across different tasks and prompts.

Maged Saeed Al-shaibani, Zaid Alyafeai, Irfan Ahmad · 0 citations
#large language models Open access Aug 2026

NormasTCU — A Brazilian Portuguese IR dataset and an evaluation of LLM-as-a-judge for relevance assessment

The results suggest that LLMs could effectively support scalable relevance assessment in specialized Portuguese corpora when evaluated using nDCG or MRR (rank-aware metrics), but they should be avoided when relying on precision or recall.

L. C. Fernandes, Marcus Vinicius Conceição de Castro, Leandro dos Santos Ribeiro et al. · 0 citations
Open access Sep 2026

Design and Implementation of a Scalable AI-Based Semantic Evaluation System for Hindi Text Using Transformer Models

Evaluating linguistically diverse descriptive answers in a consistent and accurate manner in modern digital education systems is a growing challenge, especially in low-resource languages like Hindi. Traditional lexical and rule-based grading systems cannot adequately reflect the meaning behind the words, negation, para...

Nirja D. Shah, Jyoti Pareek · 0 citations
Conference Jul 2026

Cross-Lingual Information Retrieval for Indonesian–Javanese Documents Using LLM-Based Reranking

Cross-lingual information retrieval (CLIR) for low-resource regional languages remains challenging due to limited annotated training data, morphological complexity, and linguistic diversity. This paper systematically investigates Indonesian–Javanese CLIR by integrating traditional sparse lexical retrieval, multilingual...

Raden Mohamad Adrian Ramadhan Hendar Wibawa, Ika Alfina, Evi Yulianti · 0 citations
Preprint Aug 2026

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study

This paper proposes a general optimization framework that combines a vocabulary pruning method with a targeted fine-tuning protocol for MNMT models, and reduces the vocabulary size from over 128,000 to approximately 10,000 tokens, enabling a 60% memory saving without any loss in performance.

A. A. Aliane, N. Semmar, H. Aliane · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.