Aug 2026· Conference on Control Technology and Applications· pp. 352-357· 0 citations· 36 references
Abstract
Mixed human-AI teams often operate with limited explicit communication under task uncertainty, so effective coordination requires quickly identifying which latent task or environment model best explains the observed interaction. This paper develops implicit task-hypothesis inference, enabling an AI agent to maintain a posterior over a finite set of candidate task models using its own experience and passively observed human state trajectories. We model collaboration as a task-uncertain cooperative Markov decision process (MDP) and derive a fully recursive Bayesian update that combines the agent’s experience with human data. Unobserved human actions are modeled and integrated using a bounded-rational, cooperative policy model. The resulting posterior can be used for a wide range of downstream tasks, including online decision rules that either act under the most likely hypothesis (maximum a posteriori) or account for uncertainty by weighting decisions across hypotheses according to the posterior. We analyze when human trajectories are informative by studying their sensitivity to human decision reliability and identifying conditions under which the human likelihood provides limited additional discrimination among hypotheses. Numerical results on grid-world rescue in maze environments with hidden environmental structures demonstrate faster hypothesis identification and higher cooperative returns than competing inference methods.
Online Hypothesis-Driven Conditional Action Model Learning (OHCAM), an online approach for learning action models with conditional and quantified effects from limited interactions with the environment, is presented.
Jeff Jewett, William Solow, Sandhya Saisubramanian· 0 citations
This work provides the first comprehensive empirical investigation of humans' internal models play a mediating role in feedback behaviour through a randomized controlled trial and shows that this relationship is invariant across visual contexts and is robust to three common feedback types.
Taha Shaheen, S. West, Yu Zhang· Proceedings of the Thirty-Fi...· 0 citations
Observed non-movement does not distinguish genuine latent belief inertia from small unexpressed updates, rounding, or other reporting processes, and these findings show that calibration analyses of repeated human–AI interaction should distinguish visible non-movement in elicited belief reports from updating conditional...
Shreyan Biswas, Alexander Erlei, U. Gadiraju· Proceedings of the 2026 ACM...· 0 citations
Cyber physical systems such as autonomous vehicles operate in highly dynamic environments where interactions with other autonomous and human agents is inevitable. Reinforcement learning (RL) is a well-established paradigm to allow agents to learn behaviors through interactions with the environment when a model of the e...
John Lewis, Josiah Mesler, Bhaskar Ramasubramanian· Conference on Control Techno...· 0 citations
In some multi-agent systems, the quantity to be optimized is not an externally specified reward but the information acquired about unknown properties of the environment as done in active sequential hypothesis testing (ASHT) problems. However, the ASHT literature tends to focus on finite single-agent problems with well-...
This work extends the mechanistic study of ICL to strategic multi-agent settings, introduces REE as a diagnostic tool for distinguishing reasoning from extrapolation, and provides a reusable framework for probing the boundaries of LLM reasoning in recursive belief tasks.
Y. Liu, Wen-Wen Li, Yi-Fan Dou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.