Jul 2026· Language Resources and Evaluation· Vol 60· 0 citations· 57 references
Computer Science
TL;DR
Three novel metric models are proposed: a prompt-based strategy utilizing Large Language Models to assess answers, an approach that adapts precision and recall concepts by segmenting answers into discrete information units, and a regression model trained on synthetic data to predict completeness and relevance scores.
Abstract
Evaluating the quality of long-form answers generated by Question Answering systems presents significant challenges. Traditional metrics, such as BLEU and ROUGE, often reduce the assessment to a single similarity score with a reference answer, failing to capture semantic and specific aspects of answer quality. This reliance on an aggregated score not only overlooks important dimensions but also depends heavily on the availability of reference answers, which may not always be practical or sufficient. Developing metrics capable of individually assessing specific criteria, particularly completeness and relevance, is crucial for identifying weaknesses and guiding improvements in these systems. To address these limitations, this paper introduces specialized metrics designed to evaluate completeness and relevance of long answers without the need for reference texts. We present a new dataset comprising long answers to instructional questions in Computer Science, annotated by human experts based on completeness and relevance. Building upon this, we propose three novel metric models: (1) a prompt-based strategy utilizing Large Language Models to assess answers, (2) an approach that adapts precision and recall concepts by segmenting answers into discrete information units, and (3) a regression model trained on synthetic data to predict completeness and relevance scores. Experimental results demonstrate that the proposed metrics closely align with human judgments and provide more detailed evaluations of completeness and relevance compared to traditional metrics. By enabling a more granular assessment, these metrics facilitate targeted refinements in QA systems, enhancing their ability to meet users’ informational needs more effectively.
A semantic correctness taxonomy is introduced that assigns open-ended answers to eight ordered classes, separating verbose-but-correct answers from those contaminated by hallucinated content and CAP (Context-Aware Precision), a reference-based metric that scores question-conditioned statements using bidirectional NLI.
Elitsa Yotkova, Violeta Kastreva, Petar Velkov et al.· 0 citations
Theory-based assessments remain one of the hardest grading challenges to automate. Students rarely express correct answers the way a model answer expects, and the keyword-matching tools built into most learning management systems penalise them for it, not because they are wrong, but because they phrased things differen...
E. Mgbeahuruike, Chris-Esezobor Ejodamen, Nelson-Nwanoneze Samuel et al.· British journal of computer,...· 0 citations
Large language models have improved rapidly on tasks with verifiable answers, such as mathematics and programming. Much less is known about their ability to reason about what we call conceptual questions: questions for which no ground truth is realistically accessible and no widely accepted resolution methodology exist...
By evaluating paper retrieval, evidence grounding, and answer accuracy separately, LitTraceQA provides a testbed for scientific QA systems that produce verifiable answers rather than unsupported summaries.
Evaluating open-ended large language model responses remains difficult because response quality depends not only on factual correctness and task completion, but also on subjective and scenario-dependent user experience factors. Existing benchmarks and automatic evaluators are effective for coarse-grained assessment, ye...
Tianyou Wang, Fei Yuan, Chia-Ju Miao et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.