Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or requi...
Kai-Xuan Fan, Kai-Tuo Feng, Tian-Shuo Peng et al.· 0 citations
While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. These limitations often stem from"Rollout Silencing"and low-quality gradient signals in standard sampling procedures. In this work, we propose...
Yi-Meng Ye, Shuang Chen, Wen-Xuan Huang et al.· 0 citations
The first capability-driven benchmark designed to evaluate proactive agents in dynamic, real-world settings is introduced, and comprehensive comparisons across both models and frameworks show how base model capabilities and agent framework designs jointly shape performance in real-world environments.
Zhekai Chen, Chengqi Duan, Kaiyue Sun et al.· arXiv.org· 1 citation
EVA-Client unifies the real-robot stages of the policy iteration loop within a single codebase, and consolidates major real-time inference strategies, synchronous and asynchronous execution, ACT-style temporal ensembling, Real-Time Chunking, and a naive-async ablation baseline, behind a single configuration surface.
He Yang, Yang Yi, Liyao Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.