Skip to content
Conference Open access

Evaluating Retrieval-Augmented Generation in Multi-Agent Frameworks: An Accuracy and Traceability Analysis

2026 · ITM Web of Conferences · 0 citations · 2 references

Abstract

Hallucinations in knowledge-intensive tasks are frequent in Large Language Models (LLM). This is alleviated by Retrieval-Augmented Generation (RAG) and multi-agent systems, which include external knowledge and tools. Nevertheless, the consistency of retrieved evidence and the tracing of the pipeline in multifaceted interactions are still important issues. This research evaluates the correlation between retrieval performance and final answer accuracy on a simplified RAG framework. Through a dense embedding technique and a FAISS-based vector index, the retrieval part is separated, and its performance is evaluated on SQuAD v2.0 dataset. The experimental findings starkly contrast: whereas the success of the token-level retrieval (BestTokenF1) continuously increases with the size of the retrieval pool, the overall F1 score decreases substantially in case the context is too long, and the Exact Match score is zero. This mismatch points to a severe drawback of retrieval-only architectures, which shows that too much retrieved context adds noise to the correct answer. This result highlights the need to incorporate a generative reader module without any doubt whatsoever to obtain concise responses.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.