Skip to content
#large language models Review Open access

A review of bias detection and fairness auditing techniques in LLMs

Sep 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 41 references
Ethics and Social Impacts of AI

TL;DR

A thorough literature review is provided to encapsulate prior research on bias identification and fairness auditing, categorizing the findings according to various stages of study and proposing a unified pipeline for dataset integration and a modular framework for bias auditing.

Abstract

The rise of large language models (LLMs) has sparked worries about inherent social biases and issues related to fairness. Earlier studies have investigated bias identification in word embeddings, interventions aimed at fairness in algorithms, and frameworks for auditing at the system level. Nonetheless, these methods remain disorganized, with variations in datasets, evaluation methods, and implementation processes. In this paper, we provide a thorough literature review to encapsulate prior research on bias identification and fairness auditing, categorizing the findings according to various stages of study. Additionally, we analyze the limitations in coverage and consistency of widely used benchmark datasets. To tackle these issues, we propose a unified pipeline for dataset integration and a modular framework for bias auditing. Recognized significant research gaps include the absence of intersectional bias modeling, a shortage of standardized evaluation metrics, and challenges in scalability for real-time auditing systems.

Read PDF

Similar papers

Conference Open access 2026

Bias and Fairness in LLM-Based Recruitment: A Systematic Review

A PRISMA 2020-guided systematic literature review draws on 82 studies selected from 493 records retrieved from Scopus and Web of Science and reveals a structural disconnect in the fairness-in-NLP and HCAI governance literature.

Asmae El Moutafail, Khalid Belkhoutout · 0 citations
Book Open access Sep 2026

Auditing Bias in AI-Based Hiring Systems: A Fairness Analysis of Nationality and Gender Discrimination

In recent years, artificial intelligence (AI) systems have become increasingly integrated into recruitment processes, particularly in early-stage candidate screening based on semantic matching between CVs and job descriptions. While embedding-based models promise efficiency and scalability, they also raise concerns reg...

Xhoana Shkajoti, Martina Ullasci, Marco Rondina et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Position: Fairness Failure in Generative Models is an Evaluation Problem

This position paper argues that fairness failures in generative models, albeit driven by multiple factors, are ultimately stemming from an evaluation problem: fairness findings are rarely comparable across papers or actionable for deployment decisions.

M. Vladimirova, Jean-Yves Franceschi, Thibaut Issenhuth · 0 citations
Conference Aug 2026

Bias Evaluation Framework in AI-Powered Resume Classification Using DistilBERT with SHAP Explainability and Automated Fairness Flagging

This paper proposes a multi-layer bias audit framework for AI-powered resume screening combining DistilBERT classification, SHAP explainability, automated fairness flagging, and locally-deployed LLaMA 2 interpretation. The framework achieved 74.8% of accuracy, 76.6% of precision, 74.8% of recall, and 74.7% of F1-Score...

Jotika Aleeshya Halim, Michella Arlene Wijaya Radika, Diana et al. · 0 citations
Review Open access 2026

A Fairness–Utility Evaluation Framework for Assessing Large Language Models (LLMs)

Large language models are increasingly used in contexts where their outputs can affect people directly, including hiring, admissions, and lending. This growing role makes it important to consider not only how well these models perform, but also whether their behavior is fair. Although many fairness metrics, bias benchm...

Samah Alhazmi · 0 citations
#artificial intelligence Preprint Sep 2026

Efficient Active Auditing of Multi-Group Fairness with Bias Probes

Over the past decade, Machine Learning (ML) has been trained under dual objectives: minimizing prediction error via Empirical Risk Minimization (ERM) while controlling unfairness bias. In practice, however, fairness-aware training often yields limited improvements over standard ERM, making reliable post hoc auditing es...

Ayoub Ajarra, Debabrota Basu · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.