Skip to content
Open access

From Calibrated Confidence to Calibrated Reliance: Human-Aware Confidence Communication for AI-Assisted Decision Making

Aug 2026 · ACM Transactions on Social Computing · 0 citations · 29 references

Abstract

As AI systems increasingly support consequential human decisions, the confidence they display shapes whether users accept or override their recommendations. A common assumption is that well-calibrated model confidence will induce well-calibrated human reliance. This paper shows that the two are distinct. It introduces calibrated reliance as a team-level property of human–AI decision making and proposes Reliance Calibration Error (RCE) as a metric for quantifying the gap between displayed confidence and realized team accuracy. Using 35,670 human–AI interactions from the HAIID benchmark and complementary evidence from the GRACE benchmark, the analysis identifies a systematic calibration–reliance gap: even near-calibrated models can produce miscalibrated reliance once confidence is interpreted through human judgment. The evidence is consistent with self-confidence moderating advice uptake, but the observational design does not identify anchoring as the unique mechanism. Because this gap arises at the interface between model confidence and human action, the paper develops a human-aware confidence communication framework that remaps displayed confidence without changing the underlying predictor. On held-out HAIID data, plug-in observed-outcome RCE indicates that subgroup-aware remapping can substantially reduce display–outcome misalignment, while model-predicted simulations show that bounded/global policies improve team MSE more reliably than aggressive subgroup remapping. Diagnostics also show that unconstrained remapping creates substantial boundary mass and that model-predicted counterfactual RCE is sensitive to aggressive display shifts. Bounded variants preserve much of the simulated decision-quality gain while avoiding 0/1 displays. A small prospective pilot is reported as a feasibility check rather than confirmatory validation. These findings suggest that confidence interfaces should be evaluated for human–AI team behavior, with explicit support, transparency, and robustness checks, rather than for model-side statistical fidelity alone. Code and materials for reproducing the analyses and using the confidence-display policies are available at https://github.com/OliverDOU776/From-Calibrated-Confidence-to-Calibrated-Reliance.

Read PDF