Generative video world models can now synthesize open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended generation does not imply full interaction: as a generated world expands, newly created content through navigation should expand what the agent can act upon, and as the agen...
Xin-Di Yang, Bao-Lu Li, Liam Lee et al.· 0 citations
Large reasoning models (LRMs) often suffer from overconfidence when expressing their uncertainty. Confidence-aware reinforcement learning (RL) offers a promising way to optimize calibration. However, it relies on on-policy rollouts and is thus constrained by the model's pre-RL confidence distribution, which we term con...
Shuo-Yuan Wang, Bei-Er Luo, Hao Zeng et al.· 0 citations
BB-EDGE (Benchmark-Weighted and Block-Factorized e-processes for Directed Graph Evaluation), a principled framework that represents an LLM leaderboard as a directed graph whose edges certify pairwise mean-performance advantages, with anytime-valid family-wise error rate (FWER) control, is proposed.
Hong-Fu Gao, Song-Xin Zhang, Ze-Jian Xie et al.· 0 citations
AtomEgo is presented, a systematic study of ego--robot co-training supported by a curated corpus of approximately 2,659 hours and a scalable data processing pipeline that reveals a simple principle: Data Scale * Alignment Quality -->Capability Gain; egocentric data can improve generalization, but their value depends on...
Di Wu, Dong-Chen Zheng, Jun-He Sheng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.