Skip to content
Open access

When less data is better: Exploratory privacy-by-design in AI-based human resource analytics based on synthetic data

Aug 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 21 references

Abstract

The use of Artificial Intelligence (AI) in Human Resource (HR)-related processes raises ethical and legal concerns, within the European context governed by the GDPR. This study presents a proof-of-concept computational investigation proposing an approach to securing HR data by using Large Language Models-LLMs (such as: Llama3.2, Mistral-Nemo, Gemma3), deployed on self-hosted infrastructure, which acts as a security middleware. The methodology includes two experiments applied to a multinational corporation synthetic dataset: (1) Automated Personally Identifiable Information (PII) redaction based on semantic context; (2) Measuring the impact of pseudonymization on LLM-generated promotion suitability evaluations. The results demonstrate that removing protected attributes using LLMs maintains overall consistency in LLM-generated evaluation scores while reducing identity-conditioned score variation. While 88% of LLM evaluations remain neutral under PII visibility, a persistent minority (12%) faces score alterations, which predictive models successfully forecasted with 93.2% accuracy. Furthermore, Bayesian dependency modelling identified geographic region as the strongest associated factor linked to score variation, with some non-European profiles receiving lower LLM-generated scores in specific experimental conditions. Thus, a decentralized architecture, processing data on-premise, represents the optimal solution to balance technological innovation with fundamental employee rights. The results provide technical validation of the data minimization principle, demonstrating that protecting employee identity is an essential step for building trustworthy AI systems. They should be interpreted as exploratory evidence requiring validation on real enterprise datasets rather than definitive empirical conclusions.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.