Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation
Results demonstrate that the proposed method outperforms the baselines, including an LLM without debiasing and previous calibration methods, and it is confirmed that scoring bias varies across LLMs, tasks, and score ranges, indicating the importance of measuring latent number bias as the case may be.