Skip to content
Open access

Benchmarking retrieval augmented generation LLMs for Arabic noise robustness

Aug 2026 · Frontiers in Big Data · Vol 9 · 1 citation · 44 references
Medicine

TL;DR

This paper systematically investigated the capabilities of state-of-the-art multilingual LLMs with regard to two essential RAG abilities, noise robustness and negative rejection, and revealed that all six LLMs were negatively affected when the noise ratio in the external documents was increased.

Abstract

Hallucination has become a serious concern in large language models (LLMs), as these models can generate useful yet incorrect or misleading information, which has led to growing research interest in retrieval-augmented generation (RAG) as a mitigation approach. RAG provides LLMs with access to external information, such as databases or documents, which can help them to answer users' questions. Currently, several benchmarks released measure RAG performance on various LLMs; however, evaluations of the noise robustness ability in Arabic are absent. In this paper, we systematically investigated the capabilities of state-of-the-art multilingual LLMs with regard to two essential RAG abilities, noise robustness and negative rejection. To accomplish this, we generated an Arabic benchmark consisting of 300 questions along with 6,196 documents. Then, we assessed the performance of six LLMs in relation to the two aforementioned RAG abilities. The results reveal that all six LLMs were negatively affected when the noise ratio in the external documents was increased. Under the highest noise settings at 80%, the best LLM performance was for Claude-4 sonnet, in which their performance decreased by only 4.67 percentage points. Furthermore, when it comes to the negative rejection task, there has been a significant impact on all six models. The best two models, Claude-4 sonnet and Llama-4, scored 90.67% and 85.67%, respectively, while smaller models, like GPT-3.5, scored 69.33%. Furthermore, our manual analysis reveals that many errors made by LLMs are attributed to over-caution behavior. LLMs often decline to respond probably due to training mechanisms designed to reduce hallucinations. Additionally, other errors occur when there is a high lexical similarity between the question and the words of noisy documents, which causes the model to rely on irrelevant content instead of the correct information.

Read PDF

Similar papers

Jul 2026

Visual Token Compression Enhances Robustness of MLLMs

This paper shows for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbreak attacks and hallucinations, and maintains strong results on general datasets like MME.

Shishen Gu, Jie-Quan Cui, Wen-Bo Hu et al. · 2 citations
Jul 2026

Benchmarking and AI-assisted human-like evaluation of retrieval-augmented generation for Arabic and English documents

An important practical reproducible framework for multilingual RAG benchmarking and insights for optimizing performance on resource-constrained devices are contributed and support the latent language hypothesis by suggesting an internal model bias toward high-resource languages.

B. J. Mohd, Khalil M. Ahmad Yousef, Salah G. Abu Ghalyon · 0 citations
Review Open access Aug 2026

Large Language Models Hallucinate and How Retrieval- Augmented Generation Mitigates It

It is concluded that RAG should be viewed as a grounding and evidence-access mechanism rather than a guarantee of hallucination-free generation, and applications of RAG are outlined.

Shyalaja L. N., Shantinath Patil, P. R. et al. · 0 citations
#natural language process... Preprint Sep 2026

RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety

Results are reported, showing an inverse relationship between benign and unsafe capability, strong evidence that baseline safety guardrails do not lead to downstream safety guarantees in the RAG case, and model-specific support for previous findings that even benign documents can lead to unsafe generation in retrieval-...

Adithiyan Rajan Indira Saravanan, Kathleen C. Fraser · 0 citations
Book Open access Jul 2026

Numerical Hallucinations in Retrieval-Augmented Generation: Detection and Analysis

These findings highlight numerical fabrication as a critical gap in current hallucination detection approaches and recommend the need for specialized, number-aware methods in RAG systems.

S. Singha Roy · 0 citations
Open access Aug 2026

An Optimization Framework for Retrieval Augmented Generation in Indonesian Educational Question Answering

A RAG optimization framework for Indonesian-language educational question answering using a Human-Computer Interaction learning corpus as a case study is developed and provides a procedure for selecting retrieval and generation settings for a given corpus.

I. K. R. Arthana, N. Gunantara, Made Sudarma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.