Reassessing demographic bias in face attribute classification: a statistically grounded multi-model evaluation on FairFace and UTKFace
Abstract
Face analysis systems are widely used in security, authentication, and public-sector applications; however, demographic bias and the statistical reliability of reported performance remain key concerns. Many studies rely on aggregate accuracy without quantifying subgroup disparities or uncertainty, potentially overstating model fairness. This study presents a statistically grounded evaluation of demographic bias in face attribute classification across three representative architectures, ResNet50, MobileNetV3, and a vision transformer (DeiT), using the FairFace and UTKFace datasets. Subgroup analysis is conducted across race and gender, incorporating disparity indices, bootstrap confidence intervals, and inferential statistical testing with effect size analysis. The evaluation uses an embedding-based nearest-neighbor approach to examine representation-level behavior consistently across models. Results show that race-based disparities are substantially larger than gender-based disparities across both datasets. On FairFace, race disparity gaps range from 0.1124 to 0.1266, while on UTKFace they increase significantly to 0.4726–0.4944, with large effect sizes (Cohen's d>1). In contrast, gender disparities remain smaller, with gaps between 0.0280 and 0.0582 on FairFace and 0.0194–0.0326 on UTKFace, and correspondingly small effect sizes (d < 0.13). Despite modest differences in overall accuracy across models, subgroup disparities remain statistically significant across all architectures. These findings emphasize the importance of subgroup-level evaluation, uncertainty quantification, and statistical validation for reliable fairness assessment in face analysis systems.