General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physic...
Zhiqin Yang, Chen-Xin Li, Xiao-Meng Hu et al.· 0 citations
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixt...
Rui-Bin Yuan, Jia-Hao Pan, Jun-Yan Jiang et al.· 1 citation
The model family shows gains in held-out scientific-code repair and across selected general-purpose benchmarks in code, reasoning, and knowledge, providing evidence of positive transfer from scientific experience to broader capabilities.
He-Jia Geng, Ze-Sen Huang, Hao-Yang Li et al.· 1 citation
BenchShield is presented, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation that grounds detection in a finite lifecycle model of an evaluation's reward-relevant events and achieves 96% accuracy in detecting reward hacking from infrastructure-side evidence.
Sheng-Han Zheng, Zong-Lin Di, Yimin Liu et al.· 0 citations
This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence.
Zhiqin Yang, Jing-Wen Fu, Yu-Han Liu et al.· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.