ExplorationBench turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds, and finds that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can s...
Ming Zhang, Zhen Xiang, Pei-Zhong Gao et al.· 0 citations
Atria Dawn Preview is introduced, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world.
NovGauge is presented, a human-anchored benchmark for fine-grained novelty assessment diagnosis, and a cascading diagnostic pipeline that verifies per-dimension correctness, evidence grounding, and logical support is proposed.
Guo-Qiang Zhang, Ke-Xin Tan, Ming Zhang et al.· 0 citations
Pera describes a persistent agent organized around perception and control components that continually perceive service-relevant signals from episodic task executions, internal context, and changes in the surrounding environment, and use these signals to construct lifecycle tasks.
Shi-Han Dou, Haoxiang Jia, Shichun Liu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.