Preprint
Aug 2026
Evidence Lock Before Commitment: A Frozen Interface Degrades LLM-as-Judge Evaluation
This work compares standard pairwise judging, structured one-call judging, two-call evidence locking, and three-call pointwise locking with Claude Sonnet 4.5 and GPT-5 to test an observable alternative: persist the evidence in one call and make it the exclusive input to the next.
Divyansh Singh
· 0 citations