Aug 2026· Frontiers in Big Data· Vol 9· 1 citation· 44 references
Medicine
TL;DR
This paper systematically investigated the capabilities of state-of-the-art multilingual LLMs with regard to two essential RAG abilities, noise robustness and negative rejection, and revealed that all six LLMs were negatively affected when the noise ratio in the external documents was increased.
Abstract
Hallucination has become a serious concern in large language models (LLMs), as these models can generate useful yet incorrect or misleading information, which has led to growing research interest in retrieval-augmented generation (RAG) as a mitigation approach. RAG provides LLMs with access to external information, such as databases or documents, which can help them to answer users' questions. Currently, several benchmarks released measure RAG performance on various LLMs; however, evaluations of the noise robustness ability in Arabic are absent. In this paper, we systematically investigated the capabilities of state-of-the-art multilingual LLMs with regard to two essential RAG abilities, noise robustness and negative rejection. To accomplish this, we generated an Arabic benchmark consisting of 300 questions along with 6,196 documents. Then, we assessed the performance of six LLMs in relation to the two aforementioned RAG abilities. The results reveal that all six LLMs were negatively affected when the noise ratio in the external documents was increased. Under the highest noise settings at 80%, the best LLM performance was for Claude-4 sonnet, in which their performance decreased by only 4.67 percentage points. Furthermore, when it comes to the negative rejection task, there has been a significant impact on all six models. The best two models, Claude-4 sonnet and Llama-4, scored 90.67% and 85.67%, respectively, while smaller models, like GPT-3.5, scored 69.33%. Furthermore, our manual analysis reveals that many errors made by LLMs are attributed to over-caution behavior. LLMs often decline to respond probably due to training mechanisms designed to reduce hallucinations. Additionally, other errors occur when there is a high lexical similarity between the question and the words of noisy documents, which causes the model to rely on irrelevant content instead of the correct information.
This paper shows for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbreak attacks and hallucinations, and maintains strong results on general datasets like MME.
Shishen Gu, Jie-Quan Cui, Wen-Bo Hu et al.· arXiv.org· 2 citations
An important practical reproducible framework for multilingual RAG benchmarking and insights for optimizing performance on resource-constrained devices are contributed and support the latent language hypothesis by suggesting an internal model bias toward high-resource languages.
B. J. Mohd, Khalil M. Ahmad Yousef, Salah G. Abu Ghalyon· Language Resources and Evalu...· 0 citations
It is concluded that RAG should be viewed as a grounding and evidence-access mechanism rather than a guarantee of hallucination-free generation, and applications of RAG are outlined.
Shyalaja L. N., Shantinath Patil, P. R. et al.· International Journal for Re...· 0 citations
Results are reported, showing an inverse relationship between benign and unsafe capability, strong evidence that baseline safety guardrails do not lead to downstream safety guarantees in the RAG case, and model-specific support for previous findings that even benign documents can lead to unsafe generation in retrieval-...
Adithiyan Rajan Indira Saravanan, Kathleen C. Fraser· 0 citations
These findings highlight numerical fabrication as a critical gap in current hallucination detection approaches and recommend the need for specialized, number-aware methods in RAG systems.
S. Singha Roy· Annual International ACM SIG...· 0 citations
A RAG optimization framework for Indonesian-language educational question answering using a Human-Computer Interaction learning corpus as a case study is developed and provides a procedure for selecting retrieval and generation settings for a given corpus.
I. K. R. Arthana, N. Gunantara, Made Sudarma et al.· International Journal of Adv...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.