Jul 2026· International Conference on the Theory of Information Retrieval· pp. 403-413· 0 citations· 35 references
Computer Science
TL;DR
AutoNuggetizer, a nugget-based framework, is employed to analyze ?
Abstract
Side-by-side comparisons that elicit human preferences are commonly used to assess the quality of large language model (LLM) output and have been applied to retrieval-augmented generation (RAG) systems. However, when applied to complex, information-seeking queries, these methods are limited by their inability to provide explanatory or diagnostic insights. As an alternative, nugget-based evaluations that decompose long-form answers into atomic facts have emerged as a promising strategy for RAG evaluation. In this work, we employ AutoNuggetizer, a nugget-based framework, to analyze ?5K Search Arena battles from LMArena by automatically generating and assigning nuggets, converting each model response into a quantitative score. We observe strong alignment between nugget-based Elo rankings and human preferences, exceeding the corresponding alignment achieved by LLM-as-a-judge with chain-of- thought (CoT) evaluation, while substantially reducing the number of preference inversions. Furthermore, we provide in-depth analyses including inversions, nugget quality, and shared-blindness effects. All our code is available at https://github.com/castorini/SxSNuggets.
This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.
Attune is presented, a mixed-initiative system for steerable LLM-powered scoring that performs pairwise comparisons across records to develop a global understanding first, and then resolves these comparisons into consistent score assignments-deriving scoring criteria and rules bottom-up in the process.
Bhavya Chopra, Meng Chen, Rebecca Dang et al.· 0 citations
In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these systems have been widely adopted by researchers and practitioners across a broad range of measurement tasks, driven by their strong performance, scalability, and cost-eff...
Khaoula Chehbouni, Melina Medjdoub, Florian Carichon et al.· 0 citations
This survey provides a comprehensive overview of recent advances in LLM-based evaluation, covering techniques, applications, and challenges across domains, with future directions emphasizing standardized protocols, uncertainty estimation, and human–AI collaboration.
M. Nadăş· Artificial Intelligence Revi...· 0 citations
A systematic analysis of existing evaluation frameworks and metrics for RAG-based Question Answering (QA) systems reveals urgent needs for hybrid evaluation frameworks, reasoning path traceability, combined pipeline assessment, and harmlessness evaluation, laying groundwork for the next generation of evaluation methodo...
It is argued that epistemic diversity is a useful evaluation dimension for settings where LLMs are used to support knowledge-intensive tasks, and a preliminary framework for conceptualizing and measuring epistemic diversity in LLMs is proposed, and operationalized in two domains.
Elisabeth Kirsten, N. Krämer, Muhammad Bilal Zafar· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.