Auditing Bias in AI-Based Hiring Systems: A Fairness Analysis of Nationality and Gender Discrimination
Abstract
In recent years, artificial intelligence (AI) systems have become increasingly integrated into recruitment processes, particularly in early-stage candidate screening based on semantic matching between CVs and job descriptions. While embedding-based models promise efficiency and scalability, they also raise concerns regarding fairness, as they may encode and propagate historical social biases present in training data. This study presents a controlled and reproducible audit of a hiring pipeline built on Sentence-BERT (SBERT) to verify disparate outcomes related to gender and nationality. Using a synthetic dataset of matched candidate profiles, we perform a counterfactual analysis across three operational levels: similarity scores, threshold-based screening decisions, and Top-K ranking outcomes. Results reveal systematic disparities in candidate visibility according to their gender and nationality. While score-level differences are small in magnitude, they have a significant impact on screening and ranking decisions, with female candidates and Italian profiles being the most disadvantaged groups. This work contributes a transparent audit framework and provides empirical evidence supporting the need for usage-oriented fairness evaluation in AI-based hiring systems.