Aug 2026· Applied Psychological Measurement· pp.
01466216261484166
· 1 citation· 3 references
Medicine
Abstract
Fairness in automated scoring is typically evaluated with a single global statistic contrasting a focal and reference group-an approach that can either mask a disparity that changes sign across the ability range, or overstate one by conflating it with genuine ability differences between groups (impact). We introduce condfair, an R package that adapts differential item functioning (DIF) logic to automated scoring: it estimates a conditional disparity function across ability levels, tests it with a wild-bootstrap omnibus procedure, decomposes bias into uniform and non-uniform components, and identifies candidate feature-level sources of a detected disparity via conditional SHAP disparity testing. Using the PERSUADE 2.0 essay corpus, we show a global measure can conceal a large, ability-concentrated gender disparity (marginal gap = 0.003; peak conditional disparity = 0.276, p = .001) while overstating an English Language Learner disparity by conflating it with impact (marginal gap = 0.280; conditional bias = 0.045).
Findings indicate that state-of-the-art LLMs can achieve a high degree of demographic neutrality; fundamental artefacts such as positional bias can nonetheless produce severely discriminatory outcomes; and bias auditing must extend beyond demographic parity to interaction artefacts and ecosystem structure.
A. Camargo, Rafaela Silva Figueiredo Camargo· Revista de Geopolítica· 0 citations
Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that rep...
Shu-Yi Fan, Boyuan Deng, Mengyu Xu et al.· 0 citations
The study examines whether exposure to AI-integrated Video Assistant Referee (VAR) technology changes fans' perceptions of the VAR officiating system across the dimensions of procedural fairness, trust, and satisfaction. Moving beyond prior scholarship that largely explores human-assisted review mechanisms, the stu...
Ryan Chen, Susmit S. Gulavani· Sport, Business and Manageme...· 0 citations
FairFund-Bench is introduced, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised, indicating that current LLMs robustly reproduce human deser...
This work defines the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and uses it to locate that boundary exactly.
Observed standardized test scores are the result of an endogenous process: students strategically allocate effort across multiple retake attempts to improve their outcomes. Because students differ in their ability to make these investments, the interaction between applicant strategy and institutional scoring rules---su...
Christine Ling, Diptangshu Sen, Juba Ziani· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.