Skip to content
Preprint

Proxy reliance in large language model decisions is uncalibrated to predictive evidence

Aug 2026 · 0 citations · 49 references
Computer Science

TL;DR

Large language models are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use, and causal proxy effects in four LLMs on a clinical-ranking task with known ground truth are measured.

Abstract

Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated with a protected group carry predictive value, so a changed decision can be discrimination or sound inference. We measure causal proxy effects in four LLMs on a clinical-ranking task with known ground truth, where the reliance the evidence warrants can be computed exactly and used as the reference. One audit signal yields three verdicts: over-reliance, warranted and under-reliance. Under neutral labels every model relies on proxies with no information. Informative proxies draw all three. Social field names push reliance down, below the reference in one model. Two findings explain this. Reliance severely undertracks the evidence, and social-label suppression is fragile, since in-context examples raise it above zero in every model. Accuracy-based evaluation detects none of this.

View source

Similar papers

Preprint Aug 2026

Status Association Does Not Reliably Predict Decision Leakage

Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on...

X. Abdullah · 0 citations
#artificial intelligence Preprint Sep 2026

Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment

Clinical decisions are prospective, but clinical language models are often evaluated on retrospective records that reveal the final diagnosis, treatment response, and outcome. Such evaluations may reward the use of future information rather than reasoning under the uncertainty present at the decision point. We introduc...

Misaki Matsuura, Sayantan Kumar, Ojas Kadam et al. · 0 citations
Review Aug 2026

Claim-Level Confidence Calibration for Reliable Decision Making with Large Language Models

Claim-level decomposition combined with post-hoc calibration reduces expected calibration error on factual questions while exposing failure modes on adversarial false-premise questions where decision-makers most need reliable uncertainty estimates.

Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang et al. · 0 citations
Preprint Aug 2026

When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

It is suggested that current LLMs often rely on heuristic arbitration strategies when integrating heterogeneous evidence, highlighting a failure mode for tool-augmented decision systems.

Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson et al. · 0 citations
Open access Aug 2026

Minimal but Conditional: Auditing Demographic Bias in Large Language Model Résumé Evaluation Across Commercial and Open-Weight Models

Large language models are increasingly used to read résumés and judge who advances in hiring, a task once reserved for people and now handed to systems whose reasoning is hard to inspect. Whether these models carry the demographic biases that have long shaped human hiring is therefore an urgent question, and the publis...

Vasileios Pavlopoulos · 0 citations
#artificial intelligence Review Sep 2026

Typed Decision Models: An Early Evidence Audit and Evaluation Checklist

Typed decision models (TDMs) return probability distributions over caller-defined options without generating text. TypeSafe released Jev, a commercial typed decision model, on 15 September 2026, and a small body of evaluation and replication work appeared within days. We review 28 papers posted between 19 and 24 Septem...

Li-Juan Tang, Yue-Meng Zheng · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.