Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency
Large Language Model judges are widely used to rank texts and text-generating systems through pairwise comparison, and their reliability is typically assessed via three proxies: position bias, transitivity, and pairwise agreement (self- or human-labeled). Because these proxies drive judge selection and benchmarking, a...