Measuring Position Bias in LLM-as-a-Judge for Hiring
Abstract
Large language models are commonly being used as automated screening techniques for the initial selection of candidates. However, their susceptibility to position bias-that is, the influence of the order in which resumes are presented on the scores assigned-remains largely unexplored in the context of recruitment. We conducted a controlled study to assess this position bias using Llama 3 70B, a locally deployed open-source model, on 219 real-world resumes related to a job posting for a QA engineer. Each resume was evaluated at five positions (1, 3, 5, 10, 15) within lists of 50 candidates, resulting in 1,097 ratings. A one-way ANOVA revealed asignificant effect of position on ratings $(\mathbf{F}(\mathbf{4, 1 0 7 2})=\mathbf{3 1. 1 2}, \mathbf{p}<\mathbf{0. 0 0 1})$. This magnitude of effect demonstrates that position bias extends to criteria-based recruitment tasks, and that list length can alter its direction. We address the implications for algorithmic fairness and recommend the mandatory implementation of CV randomization as well as continuous bias auditing using local models. Our fully reproducible local deployment ensures compliance with data protection regulations.