Skip to content
Review

Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent

May 2026 · arXiv.org · Vol abs/2607.18257 · 0 citations · 37 references
Computer Science

TL;DR

In a controlled study, 20 university students completed five common daily tasks using OpenClaw, a general-purpose AI agent, across tasks chosen to vary in privacy, stakes, and reversibility, and found delegation regret appeared consistently when the agent executed actions without preview, even when the output was rated as successful.

Abstract

When AI agents shift from answering questions to taking actions, users face a new problem: deciding what to delegate, to a system whose action space they cannot fully anticipate. We call the resulting dissatisfaction delegation regret, a pattern in which users regret not that the agent erred, but that it acted beyond what they would have authorized. In a controlled study, 20 university students completed five common daily tasks using OpenClaw, a general-purpose AI agent, across tasks chosen to vary in privacy, stakes, and reversibility. For each task we measured trust, perceived control, transparency, supervision burden, and approval preference on 5-point Likert scales, and collected free-text reflections analyzed through thematic coding. Three findings emerged. First, participants calibrated trust per task rather than per agent: they granted wide autonomy for advisory and low-stakes tasks but demanded confirmation for irreversible, externally visible actions. Second, irreversibility combined with external visibility, rather than stakes alone, appeared to drive trust withdrawal: the moderate-stakes email task triggered the sharpest drop in trust (M = 3.10) and the highest demand for approval (M = 4.65), whereas a high-stakes but verifiable task did not produce the same response. Third, delegation regret appeared consistently when the agent executed actions without preview, even when the output was rated as successful. We discuss implications for agent designs that expose action boundaries, support per-task autonomy policies, and separate advisory output from agentic execution.

View source

Similar papers

#human-computer interacti... Preprint Sep 2026

Trust by Design: Trust Calibration Through Non-Advisory Socratic Dialogue in Conversational Agents

As conversational AI systems increasingly operate in sensitive domains, the central challenge shifts from usability to trust calibration, ensuring that users rely on systems neither too much nor too little. Systems that provide advice or interpretations risk encouraging inappropriate reliance, particularly when users p...

R. Hassan, Nahla Aboromi, Naomi Unkelos-Shpigel · 0 citations
#machine learning Preprint Sep 2026

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

Comparing and selecting task-oriented LLM agents increasingly relies on a low-cost offline evaluation gate: persona-driven LLM user-simulators converse with each candidate, an LLM-as-a-judge scores the transcripts, and the higher-scoring agent is promoted. We introduce GAUGE, a reusable offline protocol that measures w...

Umesh Bodhwani, Thanh Tran, Kai-Lin Wei · 2 citations
Open access Sep 2026

Balancing Trust and Deliberation in Human–AI Decision Support: The Effects of Explainable AI and Cognitive Forcing Functions

Reliance on AI systems for decision support is expanding into domains where mistakes carry real consequences, which makes it important that users can weigh AI suggestions against their own judgment rather than deferring to them by default. Prior work has found that human-AI teams sometimes underperform AI alone, a patt...

Oliver Henderson · 0 citations
#artificial intelligence Preprint Sep 2026

When Intelligence Becomes Agency: A Theory of Governed, Proactive Agency for Symbiotic AI Systems

Persistent AI assistants are intended to extend human attention, memory, and coordination across changing digital and physical environments. To be truly useful they must do more than just act when asked. They must decide on their own whether a situation warrants behavior at all, when it does and in what mode, whether t...

J. Ferreira · 0 citations
Preprint Aug 2026

Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives

A knowledge-verified benchmark that first confirms through a neutral probe that an agent knows a user's entitlement, and then evaluates whether it makes false claims once an incentive to deny that entitlement is introduced, which reduces the confound between lying and not knowing and enables more rigorous auditing and...

Zhe-Yuan Liu, Wei-Liang Zhao, Xiangchi Yuan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.