This work presents $\Phi$-Bench, a benchmark for systematically evaluating LLMs on engineering the LLM infrastructure stack and spans tasks of varying complexity, ranging from localized kernel-level function completion to long-horizon implementation and end-to-end system optimization.
Lei-Lei Ding, Shu-Min Wang, Yu-Ting Huang et al.· 1 citation
The experimental results validate the effectiveness of the LabDex task design and demonstration data, and show that the benchmark supports the training and systematic evaluation of existing robotic policies across laboratory dexterous manipulation tasks at different levels, providing a foundation for further research a...
Zhipeng Tang, Sihan Chen, Sha Zhang et al.· 0 citations
Laboratory automation has made remarkable progress through robotic platforms and AI-driven scientific reasoning. However, many laboratory operations (e.g., solid--solid transfer) remain inherently dynamic and require real-time adaptation to different materials and experimental conditions. Such precision-critical manipu...
Yuhan Wu, Zhao Jin, Tao Li et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.