Skip to content
Open access

Evaluating Retrieval-Augmented Generation for personal collections: architecture, models and criteria

Sep 2026 · JLIS.it · 0 citations · 24 references

Abstract

This article presents an ongoing experiment on the use of Retrieval-Augmented Generation (RAG) architectures in library contexts, considering them not as straightforward extensions of information retrieval techniques but as document mediation devices that require explicit methodological reflection. The contribution focuses in particular on the design and analysis of an integrated evaluation framework, conceived as a structural component of system development rather than as an ex post performance check. Using the personal archive of Emanuele Artom as a case study, characterized by bibliographic and archival heterogeneity, the article examines how the quality of generated responses emerges from the interaction between knowledge base modeling, retrieval strategies and evaluation criteria. The proposed framework combines automatic metrics, LLM-as-a-judge approaches, and human-in-the-loop processes, treating divergences between automated evaluation and expert judgment as diagnostic tools for analyzing system behaviour. Preliminary results indicate that core categories of document mediation, such as relevance, completeness, citability and transparency, cannot be fully reduced to computational parameters but, instead, require continuous negotiation between automated models and disciplinary expertise. From this perspective, RAG is framed as an epistemically unstable research object, whose reliability and governability depend on the robustness and reflexivity of the evaluation processes embedded in its development.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.