$\Phi$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
This work presents $\Phi$-Bench, a benchmark for systematically evaluating LLMs on engineering the LLM infrastructure stack and spans tasks of varying complexity, ranging from localized kernel-level function completion to long-horizon implementation and end-to-end system optimization.
Lei-Lei Ding, Shu-Min Wang, Yu-Ting Huang et al.
· 1 citation