Skip to content
Open access

Beyond accuracy: completeness and relevance metrics for evaluating the quality of long answers

Jul 2026 · Language Resources and Evaluation · Vol 60 · 0 citations · 57 references
Computer Science

TL;DR

Three novel metric models are proposed: a prompt-based strategy utilizing Large Language Models to assess answers, an approach that adapts precision and recall concepts by segmenting answers into discrete information units, and a regression model trained on synthetic data to predict completeness and relevance scores.

Abstract

Evaluating the quality of long-form answers generated by Question Answering systems presents significant challenges. Traditional metrics, such as BLEU and ROUGE, often reduce the assessment to a single similarity score with a reference answer, failing to capture semantic and specific aspects of answer quality. This reliance on an aggregated score not only overlooks important dimensions but also depends heavily on the availability of reference answers, which may not always be practical or sufficient. Developing metrics capable of individually assessing specific criteria, particularly completeness and relevance, is crucial for identifying weaknesses and guiding improvements in these systems. To address these limitations, this paper introduces specialized metrics designed to evaluate completeness and relevance of long answers without the need for reference texts. We present a new dataset comprising long answers to instructional questions in Computer Science, annotated by human experts based on completeness and relevance. Building upon this, we propose three novel metric models: (1) a prompt-based strategy utilizing Large Language Models to assess answers, (2) an approach that adapts precision and recall concepts by segmenting answers into discrete information units, and (3) a regression model trained on synthetic data to predict completeness and relevance scores. Experimental results demonstrate that the proposed metrics closely align with human judgments and provide more detailed evaluations of completeness and relevance compared to traditional metrics. By enabling a more granular assessment, these metrics facilitate targeted refinements in QA systems, enhancing their ability to meet users’ informational needs more effectively.

Read PDF

Similar papers

#natural language process... Preprint Sep 2026

How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation

A semantic correctness taxonomy is introduced that assigns open-ended answers to eight ordered classes, separating verbose-but-correct answers from those contaminated by hallucinated content and CAP (Context-Aware Precision), a reference-based metric that scores question-conditioned statements using bidirectional NLI.

Elitsa Yotkova, Violeta Kastreva, Petar Velkov et al. · 0 citations
Open access Sep 2026

Design and Implementation of a Sentence-BERT-Driven Semantic Matching Model for Automated Evaluation of Theory Answers

Theory-based assessments remain one of the hardest grading challenges to automate. Students rarely express correct answers the way a model answer expects, and the keyword-matching tools built into most learning management systems penalise them for it, not because they are wrong, but because they phrased things differen...

E. Mgbeahuruike, Chris-Esezobor Ejodamen, Nelson-Nwanoneze Samuel et al. · 0 citations
Jul 2026

A dataset of rated conceptual arguments

Large language models have improved rapidly on tasks with verifiable answers, such as mathematics and programming. Much less is known about their ability to reason about what we call conceptual questions: questions for which no ground truth is realistically accessible and no widely accepted resolution methodology exist...

Emery Cooper, Caspar Oesterheld, Linh Nguyen et al. · 0 citations
Conference Jul 2026

AutoQABench: A Three-Level UX Benchmark for Automated Evaluation of Open-Ended LLM Responses

Evaluating open-ended large language model responses remains difficult because response quality depends not only on factual correctness and task completion, but also on subjective and scenario-dependent user experience factors. Existing benchmarks and automatic evaluators are effective for coarse-grained assessment, ye...

Tianyou Wang, Fei Yuan, Chia-Ju Miao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.