Aug 2026· Proceedings of the 2026 ACM Conference on Human-AI Complementarity and Alignment· 0 citations· 65 references
Computer Science
TL;DR
Observed non-movement does not distinguish genuine latent belief inertia from small unexpressed updates, rounding, or other reporting processes, and these findings show that calibration analyses of repeated human–AI interaction should distinguish visible non-movement in elicited belief reports from updating conditional on movement.
Abstract
Repeated human–AI interaction is often analyzed through pooled belief-updating slopes: users observe AI successes and failures, revise reported beliefs in the feedback-consistent direction, but appear conservative on average. We show that such averages can obscure an important distinction between whether an elicited belief report changes at all and how it changes conditional on movement. We refer to this measurement-aware decomposition as the belief update gate. Reanalyzing a multi-task human–AI decision-making dataset with 240 participants, 7,200 trials, and three task domains, we find substantial non-movement in reported beliefs: 67.3% of trial-level belief changes are exactly zero, and 76.4% are smaller than five percentage points. Separating non-moving from moving reports changes the descriptive interpretation of pooled conservatism: the within-trajectory slope rises from 0.494 overall to 0.949 among rows with nonzero movement. Since this latter estimate conditions on observed movement, we interpret it as a descriptive decomposition rather than as evidence of a near-Bayesian latent learning process. Complementary hurdle style analyses (i.e., modeling zero vs. non-zero changes before predicting update magnitude) show that the absolute discrepancy between feedback and entering belief predicts whether a report changes, while the signed feedback discrepancy predicts the direction and magnitude of change among reports that move. Importantly, observed non-movement does not distinguish genuine latent belief inertia from small unexpressed updates, rounding, or other reporting processes. These findings show that calibration analyses of repeated human–AI interaction should distinguish visible non-movement in elicited belief reports from updating conditional on movement rather than treating reported beliefs as a single continuous updating process.
This work provides the first comprehensive empirical investigation of humans' internal models play a mediating role in feedback behaviour through a randomized controlled trial and shows that this relationship is invariant across visual contexts and is robust to three common feedback types.
Taha Shaheen, S. West, Yu Zhang· Proceedings of the Thirty-Fi...· 0 citations
LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants...
Sebastian Pohl, Harsh Mehta, Pranav Mambayil et al.· arXiv.org· 0 citations
Across simulations, a human-subject study, and a real-world GEM vehicle deployments in repeated lane-merging scenarios, BAIT achieves task performance comparable to baselines that optimize long-term influence through unpredictability while yielding significantly higher user trust.
Simulations show that over-reliance on a weak AI is especially harmful, and that diversifying AI signals across users can better keep the crowd informative, and conclude with implications for understanding human-AI interaction in information spread and designing misinformation interventions.
Zhuoran Lu, Weilong Wang, Yang-Yang Yu et al.· Proceedings of the 2026 ACM...· 0 citations
A novel Belief-Guided architecture that disentangles the Policy head from a distinct Belief head is introduced, enabling professional-level play on limited hardware where massive MCTS is infeasible.
Mehrad Yaghoubi, A. Bastanfard, Abbas Jalilvand et al.· arXiv.org· 0 citations
This work proposes R$^2$-MAD (Remember and Reweight for Multi-Agent Debate), a framework that equips agents with an experience memory accumulated from past debates that achieves consistent improvements over existing single-agent and MAD baselines.
Xuan-Fa Jin, Zhijian Ma, Yong-Cheng Zeng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.