These results establish scope matching as a complementary control for persistent agent memory: certification determines whether an edit is supported, while retrieval scope determines where that evidence authorizes its use.
Ye-Zhou Cheng, Run-Jia Du, Ze-Ming Liu et al.· 0 citations
Real-world GUI usage frequently involves workflows that span multiple devices and platforms, requiring the transfer of intermediate results, maintenance of shared state, and coordination across heterogeneous environments. However, existing GUI benchmarks overwhelmingly evaluate agents on single-device, statically defin...
Zi-Xiang Chen, Yu-Heng Lu, Zihao Cheng et al.· 0 citations
VGEBench is introduced, a comprehensive benchmark designed to evaluate the generalizable visually grounded exploration capabilities of VLMs, and a Logic-Driven State Machine framework is constructed, which simulates multi-turn interaction loops, compelling agents to achieve goals by active visual perception and feedbac...
Linshujie Zheng, Zeming Liu, W. Chen et al.· 0 citations
A new task, Behavior-Aware Travel Planning, which infers user preferences directly from past behaviors and generates personalized travel plans and proposes B2T-Agent, a reinforcement learning-based agent that leverages user behavior trajectories, interacts with external tools for preference-aligned retrieval, and maint...
Zihao Cheng, Yingyu Shan, Hongru Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.