Skip to content
Preprint

Evidence Lock Before Commitment: A Frozen Interface Degrades LLM-as-Judge Evaluation

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

This work compares standard pairwise judging, structured one-call judging, two-call evidence locking, and three-call pointwise locking with Claude Sonnet 4.5 and GPT-5 to test an observable alternative: persist the evidence in one call and make it the exclusive input to the next.

Abstract

LLM judges are often asked to extract criteria and evidence before choosing between candidate answers. This workflow assumes that the intermediate record preserves the information needed for a later verdict. For reasoning-capable models, visible field order does not reveal internal decision order, so we test an observable alternative: persist the evidence in one call and make it the exclusive input to the next. Across 24,000 judgments over HelpSteer3, FeedbackQA, and CoVal, we compare standard pairwise judging, structured one-call judging, two-call evidence locking, and three-call pointwise locking with Claude Sonnet 4.5 and GPT-5. Evidence locking reduces agreement with released human preferences by 4 to 6 percentage points and increases answer-order inconsistency by 8 to 10 points relative to structured one-call judging. Pointwise locking is also harmful, while structured evidence elicitation remains close to standard judging. The result holds for both judges and all three datasets. Persisted evidence can support auditability, but it should not replace the source answers at decision time.

View source

Similar papers

Preprint Aug 2026

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines

An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships. We show that the cost of that decision depends less on judge accuracy than on the decision rule the judge is embedded in. On frozen candidate pools from four GRPO policies, an unconstrained scalar DeepSeek-...

Yi-Yao Zhang, Diksha Goel, Hussain Ahmad et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators

Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated response, judge instructions, and scoring rubric. Existing systems examine important subsets of these failure modes, but auditing a...

Jackson Hassell, Farima Fatahi Bayat, Pouya Pezeshkpour et al. · 0 citations
Preprint Aug 2026

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.

V. Rodionov, Shamil Assylbekov · 0 citations
#artificial intelligence Preprint Aug 2026

Commit-first LLM judging inherits the judge's own errors

commit-first judging does not remove the anchor that gets gamed, it moves it from the candidate to the judge's own answer, so evaluation is only as good as the judge is at the task, so evaluation is only as good as the judge is at the task.

Idil Gozel · 0 citations
Conference 2026

When Verification Hurts: The Cost of Overriding Abstention in Two-Stage Web Agents

This study cautions against transplanting verification into grounding pipelines and identifies calibrated abstention as a property worth preserving and proposes an abstention-aware verifier that intervenes only under sufficient candidate coverage and confidence.

Duchen Li · 0 citations
#natural language process... Preprint Sep 2026

Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal

A natural way to cut reasoning-model inference cost is to repeatedly probe a single partial trajectory for its current answer and stop once probes agree -- self-consensus. We ask whether any such rule is both safe and token-saving, and whether one can be selected once and reused. A preregistered sweep of 3,520 consensu...

Yun-Xian Mo, Dong-Hao Zhao, He-Jia Geng · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.