Skip to content
Open access

Evaluating performance bias in face-to-BMI vision transformer models across diverse human populations

Sep 2026 · bioRxiv · 0 citations · 100 references
Biology

Abstract

Computer vision models that estimate body mass index (BMI) from facial features offer a non-invasive, low-cost alternative to physical measurement, with uses in telemedicine, emergency care where a scale or measuring tools aren’t available, automated self-monitoring, and large-scale epidemiological research. Most of these models, however, are trained on government records, social media images, and celebrity photographs, sources that introduce dataset biases and fail to represent the general public. This study tests how well a face-to-BMI machine learning model generalizes across populations, specifically how morphological diversity and population-specific training data affect cross-cultural accuracy. We trained and evaluated Vision Transformer (ViT-H/14) models on paired BMI measurements and facial photographs from four Indigenous populations: the Orang Asli of Malaysia, the Ju/’hoansi of Southern Africa, the Sama residing in the Philippines, and the Tsimane of Bolivia. To evaluate how training data composition affects predictions, we compared four training strategies, from single-population models (focal models) to models trained on the full combined global dataset (global models). In-distribution training always produced the best performance. Models exposed to a target population’s morphology, whether focal or global, consistently predicted BMI most accurately for that population. But when a target population differed from the training sample, adding more cross-cultural variation to training improved out-of-distribution predictions. Therefore, training on a population’s own data works best when that data exists, and training on data spanning a wide range of human morphology is the strongest fallback when it doesn’t. These findings suggest that while target population training data produces the most accurate results, training on datasets that capture global morphological variation substantially improves performance in unrepresented populations. Broader diversity in training data is essential for developing machine learning health tools that generalize reliably across human populations. Author summary Computer vision models that estimate body mass index (BMI) from photographs offer a promising, non-invasive, and low-cost tool for telemedicine, emergency care, and large-scale global health research. However, the majority of published models rely on a limited number of training datasets, all drawn from urbanized, market-integrated populations with measurements that are self-reported, or estimated. While limited studies highlight risks for ethnic bias in current facial analysis models, it remains unknown whether face-to-BMI models fail to generalize across diverse global populations. To test for this, we trained a Vision Transformer model using paired photographs and BMI measurements collected by anthropologists in four morphologically and geographically distinct Indigenous populations. We found that these models are most accurate when trained on the same populations they will be assessing. When that specific data is unavailable, training on a broad cross-cultural dataset can provide a strong alternative. Our findings indicate that diverse training datasets are essential for developing accurate face-to-BMI models.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.