Jun 2026· arXiv.org· Vol abs/2606.06081· 0 citations· 44 references
Computer Science
TL;DR
This paper develops the first formal framework for measuring appropriate reliance on set-valued AI advice within the sequential judge-advisor paradigm, spanning both classification and regression tasks.
Abstract
Appropriate reliance on AI advice has become a central research theme in human-AI collaboration. Existing frameworks have focused exclusively on point predictions as AI advice. However, set-valued AI advice (e.g., discrete sets or continuous intervals) is increasingly being used to communicate uncertainty and improve human decision making. In this paper, we develop the first formal framework for measuring appropriate reliance on set-valued AI advice within the sequential judge-advisor paradigm, spanning both classification and regression tasks. For classification, we first introduce the dimensions that are necessary for evaluating set-valued AI advice. We then define two metrics: correct reliance rate on AI and correct reliance rate on self, which jointly characterize appropriate reliance in this setting. For regression, we introduce quantity of AI reliance and quality of AI reliance, which respectively measure whether a decision maker utilized the AI advice and whether their reliance helped them get closer to the ground truth relative to their initial estimate. Through the application of our framework, we demonstrate how these metrics capture important nuances in human-AI collaboration that existing measures overlook.
Findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues, which position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance.
A. Kapadia, Eshwar Chandrasekharan, Koustuv Saha· 0 citations
While participants rated hedged and unhedged AI as equally trustworthy and likely to be correct, they were significantly less likely to follow hedged advice in a binary choice, and how linguistic markers can be used to calibrate user reliance to model certainty is discussed.
Laura Spillner, Johanna Rockstroh, Nina Wenig et al.· International Conference on...· 0 citations
This work develops a six step BBN framework and illustrates it to model customer intention to consult a doctor in an alternative healthcare system and reveals that while self efficacy appears to be a major factor, its actual causal impact is small.
Kumar Rahul, Shovan Chowdhury Indian Institute of Management Kozhikode, Kerala et al.· 0 citations
It is argued that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.
M. Economides, Paul M. Sacher, Samuel Salzer et al.· 0 citations
A practical decision framework for four foundational measures - Entropy, KL divergence/cross-entropy, Mutual Information, and Transfer Entropy is provided, organized around three prescriptive questions for each: what question does the measure answer and in which AI context; which estimator is appropriate for the data type and dimensionality; and what is the most dangerous misuse.
Nikolaos Al.Papadopoulos, Konstantinos E. Psannis· 0 citations
This work experimentally explores the question: is it always optimal to train the assessor for the target metric, or could it be better to train for a different metric and then map predictions back to the target metric?
Daniel Romero-Alvarado, Fernando Mart'inez-Plumed, José Hernández-Orallo· Machine-mediated learning· 0 citations
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.
MIT News · Artificial Intelligence· news.mit.eduJul 7, 2026
The professor of physics and inaugural director of the NSF AI Institute for Artificial Intelligence and Fundamental Interactions will lead LNS and continue his research in particle physics.