Skip to content

LLM Judges Can Be Too Generous When There Is No Reference Answer

Jul 2026 · arXiv.org · Vol abs/2607.12885 · 1 citation · 50 references
Computer Science

TL;DR

The results emphasize the need for calibrating the LLM judges with a sample with reference-aware evaluation before using them in reference-free setups reliably, and the methodology provides a blueprint for researchers and practitioners in doing such calibration of LLM judges for other tasks.

Abstract

LLM judges are increasingly being used to evaluate open-ended model responses, often in no-reference settings where a ground-truth answer is unavailable. However, can they reliably assess in such evaluation setups? We explore this question in this paper through a two stage pipeline with a) calibration experiments that assess the judge model's knowledge of the task it is evaluating, and b) sensitivity experiments that assess how the judge model's performance is impacted by the presence and positioning of the reference answer in the prompt. Across experiments covering three languages, we show that the judge models we evaluated tend to over-credit incorrect answers in the absence of a reference answer, and adding reference answer information to the prompt flips the judge model's correct/incorrect decisions by as much as 85% in some experimental settings. Comparison with a subset of human annotations shows that these reference-driven changes generally align with human judgments. Our results emphasize the need for calibrating the LLM judges with a sample with reference-aware evaluation before using them in reference-free setups reliably, and our methodology provides a blueprint for researchers and practitioners in doing such calibration of LLM judges for other tasks.

View source

Similar papers

#natural language process... Preprint Sep 2026

When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings

Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faster and cheaper than human evaluation. However, a judge may produce consistent scores without necessarily agreeing with human evaluators. In this work, we study this issue...

Aakash Kumar Tiwari · 1 citation
Preprint Aug 2026

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

A risk-controlled framework that calibrates uncertainty thresholds on a held-out set so that the false discovery rate among accepted verdicts remains below a user-specified level~$\alpha$ with high probability, using finite-sample Clopper--Pearson intervals is proposed.

Sher Badshah, Ali Emami, Hassan Sajjad · 1 citation
Preprint Aug 2026

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

It is shown that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values, and effective mitigation must be validated for the intended model and task or domain.

A. Kapetanović, Kemal Altwlkany, Andro Merćep et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has studied biases in pairwise LLM-as-a-judge settings; in this paper, we focus on absolute scoring tasks, which mirror more realistic use case...

Gemma Zhang, Prachi Badarayani, Asmi Kumar et al. · 0 citations
#artificial intelligence Preprint Aug 2026

When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators

Testing an LLM on two e-commerce data quality tasks, entity matching and brand mislabeling, against rule based baselines and human verified ground truth, under both zero-shot and few-shot prompting suggests that the value of using an LLM over traditional methods depends heavily on the task.

Praphulla Lal Shrestha · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.