Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generation, making them attractive for latency-sensitive applications. GUI agents represent a natural testbed for this paradigm, as they must repeatedly perceive screen states and emit structured, spatially...
Zhangxuan Gu, Hao-Xin Chen, Qiujieli Qin et al.· 0 citations
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong...
Chu-Yan Chen, Hao-Xin Chen, Kun Chen et al.· 1 citation
UI-Venus-2 is presented, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework that integrates safety-aware mechanisms to ensure controlled execution of consequential actions.
Venus Team, Zhuo-Hang Cai, Hao-Xin Chen et al.· 1 citation
MAGA is introduced that re-allocates training signal according to the structured action and suppresses unnecessary or invalid distillation signals and focuses learning on erroneous actions.
Hang Yan, Zhangxuan Gu, Bei-Tong Zhou et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.