Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· pp. 7346-7354· 0 citations· 37 references
TL;DR
A large-scale audit of sociodemographic error disparities in both lexical- and transformer-based models of county-level health outcomes, over a dataset cover billions of community-mapped messages highlights a critical accuracy–fairness trade-off in community-level models for public health tasks.
Abstract
Many high-stakes social applications of AI, such as public health surveillance and policy planning, operate at the community- rather than individual-level. However, most AI fairness research operates at the individual- or data-level (i.e. document or image) and rely on metrics defined over discrete demographic categories rather than population-level demographic proportions.
In this work, we first introduce the Bilateral Concentration Index (BCI) to quantify nonmonotonic error disparities missed by the category-based metrics used at individual or data-levels.
Then we conduct a large-scale audit of sociodemographic error disparities in both lexical- and transformer-based models of county-level health outcomes, over a dataset cover billions of community-mapped messages. While all tasks had significant disparity, the size varied widely depending on the outcome and model, from BCI of 2.1% for predicting life satisfaction to 17.0% for predicting fair or poor health. We further evaluate four approaches for incorporating sociodemographic information, as potential bias mitigation strategies, finding that while demographic inclusion consistently improved predictive accuracy, it frequently amplified error disparities. The largest disparities were associated with education and income (BCI = 2.7–16.4%), often reducing accuracy for low-income, and in some cases, high-income communities. These findings highlight a critical accuracy–fairness trade-off in community-level models for public health tasks, demonstrating how seemingly beneficial modeling choices can lead to increased disparities which could disadvantage communities if model predictions are used to make policy decisions in geographic areas where this public health data is not available.
It is shown that requiring sufficient individual-level SDoH survey data results in significant selection bias and sample reduction in AoU, and that area-level SDoH metrics contribute to disease prediction independently of individual-level measures.
M. Hysong, A. Manning, Michael D. Green et al.· Communications Health· 0 citations
It is established that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard, and subgroup-disaggregated reporting as a default standard for personalized configurations.
Junjie Luo, Xuzhe Zhi, Rui Han et al.· 0 citations
India lacks a standardised, individual-level health risk metric equivalent to the CIBIL credit score, leaving insurers to price risk largely on self-declaration and occasional point-in-time screening. This paper constructs a prototype health risk score for the Indian market — as a diagnostic exercise rather than a commercial product — to identify the data gap and quantify the cost of the data gap. We compare a Full Clinical Model incorporating laboratory biomarkers against a Restricted Predictor Model without them; the clinical model achieves a PR-AUC 33% higher, with six of the ten most predictive variables originating from laboratory results. Yet in practice fewer than 5% of policyholders undergo screening at underwriting, and existing medical records remain largely un-digitised and fragmented. We argue that this data deficit sustains a self-reinforcing "Market-Failure Feedback Loop": pervasive information asymmetry prevents refined risk pricing, forces an implicit cross-subsidy from healthy to less-healthy policyholders, suppresses low-risk participation, and surfaces as claim denials that deepen a public trust deficit and entrench under-penetration. Drawing on the precedent of credit bureaus in Indian lending, the paper contends that closing this gap requires a government mandate for standardised health-data reporting. It further examines the central risk such a mandate creates — exclusion of high-risk individuals — and outlines protections (risk pooling, policy design, constrained use, and universal-coverage backstops) intended to ensure that fuller disclosure leads to fair pricing rather than denial of coverage.
Agastya Bhandari· International Journal of Soc...· 0 citations
The unfairness tree (utree) is proposed, a data-driven recursive partitioning framework for identifying subgroups with differential model performance that exhibits nominal empirical type I error rates and good ability to detect, quantify, and characterize performance discrepancies defined by higher-order variable interactions.
Decisions about healthcare funding and delivery in Aotearoa New Zealand, as well as the monitoring of health service performance and outcomes, are driven by readily available data, in particular from administrative health datasets. Most of these national health data collections are generated through the delivery of secondary health services. Further, apart from hospitalisation and mortality collections, these lack diagnostic coding, rendering the burden and cost of chronic health conditions predominantly managed in primary care largely invisible. Many of these types of health conditions disproportionately affect women, who, despite their longer life expectancy, spend 25% more time in poor health than men, according to international research on the women's health gap. This gap is driven by conditions occurring only in women (e.g., premenstrual syndrome, endometriosis, polyendocrine metabolic ovarian syndrome) or with higher burden in women (e.g., anxiety, depression, migraine). Addressing this gap could add US$1 trillion to the global economy. In Aotearoa New Zealand, major improvements in national health data collections are urgently needed to assess the cost of the women's health gap and the burden of chronic diseases that have high social and economic impact but are undetectable or difficult to survey in our existing administrative datasets. We illustrate these issues using the example of migraine disease, the most disabling neurological condition in Australasia that also affects at least twice as many women as men.
Fiona Imlach, Natalia Boven, Vanessa Selak· The New Zealand medical jour...· 0 citations
Background Large language models (LLMs) can significantly broaden access to physical-activity guidance. However, advice that implicitly assumes available financial resources, reliable transportation, specialized equipment, or local facilities can be difficult for individuals in resource-constrained environments to act upon. We evaluated whether such resource assumptions systematically vary across socioeconomic settings when underlying health needs remain fixed. Methods We developed a county-aware, matched-counterfactual auditing framework integrating U.S. public health and socioeconomic data, synthetic patient profiles, and structured LLM outputs. To evaluate resource burden, we constructed a transparent access-cost proxy that renders each coded component directly inspectable. We evaluated accessibility-aware prompting, output reranking, alternative weighting schemes, and a bounded three-model comparative panel (including DeepSeek) to audit and mitigate socioeconomically driven bias. Results In the full DeepSeek model panel, accessibility-aware prompting reduced the access-cost proxy gap between high- and low-socioeconomic status (SES) counties from 0.306 to 0.139. Subsequent output reranking maintained this narrowed gap while simultaneously improving coarse alignment with physical-activity volume and intensity guidelines. Sensitivity analyses using alternative weighting schemes preserved the overall comparative ordering, while the three-model comparison demonstrated model-specific variations in resource-assumption responses. Discussion This study establishes a reproducible framework that connects equity-oriented health-recommender principles to traceable output auditing and targeted bias mitigation. By offering a transparent approach to evaluating implicit resource assumptions, this work provides researchers and practitioners with an actionable foundation for downstream expert, user, and implementation validation.
He-Gui Bao, Tian-Wei Yu, Jun-Yue Wang et al.· Frontiers in Public Health· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.