Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primary responses. The design separates forced latent association from matched consequential decisions across academic selection, professional hiring, research fellowship selection, and legal-aid intake. Elite-coded surnames received higher forced high-status probability mass than common surnames in seven of eight models and higher mass than rare-frequency controls in all eight. Yet elite-minus-common decision effects were close to zero for most systems. Five models were statistically equivalent within a predeclared (Plus-Minus)0.10 standard-deviation margin, while the remaining three were imprecise or borderline, with no consistent elite advantage. Association strength did not reliably predict decision leakage across models (r = 0.201, p = 0.633) or across frozen surname-pair-by-model cells (r = 0.065, p = 0.565). The central result is a measurement dissociation: latent social association and consequential treatment are empirically distinct constructs. Evaluations should measure the transition from association to action directly.
Large language models are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use, and causal proxy effects in four LLMs on a clinical-ranking task with known ground truth are measured.
Can a restricted computational model predict well while omitting distinctions required by its explanatory task? We define boundary sufficiency relative to an outcome, a representation, and a declared family of input interventions. Exact sufficiency requires a common response law on every fiber of the retained represent...
Ranks depend on the observations used for comparison. Reusing those observations can add association between forecast and outcome rank contrasts even when the evaluated forecast and outcome stay fixed. We characterize assignments that preserve association between specified population-rank contrasts, including forecast...
In selection processes, decision-makers sometimes rely on selection tests that are biased by irrelevant attributes such as gender, race, or appearance. If left uncorrected, this can lead to suboptimal decisions and discrimination. Decision-makers can mitigate this bias by treating the irrelevant attributes negatively a...
Hagai Rabinovitch, D. Budescu, Yoella Bereby-Meyer· Collabra: Psychology· 0 citations
Rare clinical outcomes pose a difficulty deeper than ordinary class imbalance: a penalized logistic model can return finite, stable-looking coefficients before the data support a reliable threshold decision. We formulate the accrual question as decision-targeted sequential certification. On a prespecified finite monito...
Comparing two populations at the same physical covariate value requires more than conditional means or isolated target-point decisions: researchers may need evidence about an entire conditional-distribution ordering over a continuum, even when covariate margins differ. This paper makes that common-value comparison esti...