Graphical User Interface (GUI) grounding is essential for autonomous agents to map natural language instructions to precise screen coordinates. However, existing supervised fine-tuning and reinforcement learning methods are constrained by the high cost of annotation, creating a scalability bottleneck. In this paper, we...
Yi-Zhou Liu, Fei Tang, Yuchen Yan et al.· 0 citations
Hierarchical vision-language-action (VLA) systems consist of a high-level vision-language planner and a low-level action expert that generates continuous actions. This hierarchical design has practical value only if the planner can generate plans fast enough to meet real-time control requirements, and the resulting pla...
Chuan-Liang Xie, Bo-Yu Ma, Gen Li et al.· 0 citations
Complex 3D spatial text to image generation requires models to convert natural language into stable visual geometry, not merely semantic appearance. Existing prompt-driven or layout-conditioned methods improve controllability, but often lack an optimizable and verifiable spatial intermediary before visual sampling. As...
Ziyun Qian, Zi-Zhi Chen, Yi-Zhou Liu et al.· 2 citations
Generative models have advanced image-conditioned 3D content creation, yet generating controllable and executable 3D scenes from a single image remains challenging. Existing 3D generative approaches can synthesize visually plausible objects and scenes, but their spatial layout estimation is coupled with specific asset...
Ding-Kang Yang, Yi-Zhou Liu, Wen-Dong Cheng et al.· 0 citations
Omni-modal models have expanded multimodal interaction across vision, audio, speech, and language. However, their training is predominantly organized around semantic descriptions and general-purpose objectives, leaving physical attributes, interaction states, and causal mechanisms only partially specified. This gap is...
Yi-Zhou Liu, Jing-Hang Han, Kai Qiu et al.· 0 citations
This work proposes RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD.
Yan Yu, Zhengxi Lu, Yi-Zhou Liu et al.· 1 citation
Automatic translation quality metrics trained on general-domain corpora systematically fail on social media content, where communicative intent is encoded in culturally loaded expressions (internet slang, homophonic ciphers, and platform-specific idioms) rather than surface token patterns. We conduct a systematic empir...
Yi-Wen Qiu, Linjuan Wu, Ding-Ming Li et al.· 0 citations
TRACER is proposed, which formulates compression as a sequential per-tool decision problem, and demonstrates the value of consequence-aware, per-tool context retention for improving the efficiency of long-horizon language agents.
Zihao Lin, Ye Wu, Mengning Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.