Skip to content

Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI Actions

Aug 2026 · Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies · 0 citations
Computer Science

TL;DR

The Act2Intention framework is proposed, which builds an active mobile agent by integrating understanding, predicting user intentions, and executing decisions, and establishes a standardized platform for developing and evaluating proactive agents and consequently paves the way for research on intention-driven human-computer interaction.

Abstract

Mobile GUI Agents powered by multimodal large language models (MLLMs) show promise in human-computer intelligence. However, current research primarily focuses on reactive task execution while lacking a comprehensive understanding-prediction-execution process for user intentions, which are the core requirements of active agents. In this paper, we propose the Act2Intention framework that builds an active mobile agent by integrating understanding, predicting user intentions, and executing decisions. First, we construct the Act2Intention Bench through data collection and validated generation, comprising 72,511 intentions and over 700,000 actions across 52 apps, thereby establishing the first benchmark for evaluating proactive agents via continuous intention-action trajectories. We further develop the Act2Intention Agent, achieving proactive services through Proactive-oriented Intention Understanding, Personalized Proactive Intention Prediction, and Experience-guided Intention Execution. Experimental results show that supervised fine-tuning on Act2Intention Bench yields absolute improvements of +32.0 Acc-S, +10.25 Acc-S, and +6.9 SSR points over non-fine-tuned counterparts under the same agent framework for intention understanding, prediction, and execution, respectively. This success underscores the necessity and value of the Act2Intention Bench, which establishes a standardized platform for developing and evaluating proactive agents and consequently paves the way for research on intention-driven human-computer interaction.

View source

Similar papers

Preprint Aug 2026

StepReflect: Structured UI Transition Reflection for Mobile GUI Agents

StepReflect is proposed, which formulates per-step GUI reflection as supervised structured prediction conditioned on explicit transition specifications and paired visual evidence, and established as a practical, locally deployable alternative to repeated frontier-model reflection for long-horizon mobile GUI agents.

Linqiang Guo, Wei Liu, Li Gu et al. · 0 citations
Preprint Sep 2026

Cognitive Action Reasoning for Proactive Robots from Human-Centered Multimodal Observations

Robots operating in human-centered environments are typically designed to execute explicit instructions, and most robot-learning datasets likewise pair observations with task instructions or low-level actions. Although recent work has begun to explore proactive embodied assistance, existing resources target different s...

Zhi-Hao Gu, Kenny Zhu, Yuan-Feng Wu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning

WM-R1 is the first reinforcement learning framework that trains mobile GUI agents with world models instead of real environments, eliminating the need for real-environment interaction, supports massively parallelized and step-level granularized trajectory generation grounded in world models, and introduces a multi-dime...

Yu Han, Tianwen Qian · 0 citations
Jul 2026

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

Qwen-UI-Agent is presented, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments, that sets state-of-the-art performance on mobile-use benchmarks while delivering competitive performance on computer- and browser-use tasks against frontier models.

Hanzhang Zhou, Panrong Tong, Xu Zhang et al. · 4 citations
Aug 2026

Nani AGI: A Personalized Multimodal AI Architecture

Nani AGI is a personalized, multimodal, and proactive artificial intelligence architecture designed to support natural and continuous human–AI interaction beyond conventional question-and-answer systems. The proposed architecture integrates large language model-based conversational reasoning with persistent memory, voi...

Devireddy Lohith Reddy · 0 citations
Preprint Aug 2026

Agentic Router: An Execution-Grounded Continual Learning Approach With Memory

An execution-grounded dual-path consequence-aware agent for CLI-based SONiC operations, which generates multiple complete actions, predicts their execution consequences, and selects the final action through utility- and risk-aware reranking is proposed.

Yuxuan Chen, Rong-Peng Li, Zhi-Feng Zhao et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.