DataClawEval is introduced, the first comprehensive benchmark designed specifically to evaluate the end-to-end task completion capabilities of autonomous agents in real-world data engineering scenarios, and it comprises 100 rigorous, end-to-end tasks spanning five execution engines.
Debin Meng, Jiaming Yang, Zefang Zong et al.· 0 citations
OptiDSL is proposed, a framework that shifts the focus from rigid MILP formulations to domain-specific language (DSL) representations, and enables seamless integration with a diverse library of specialized solvers, ranging from traditional heuristics to modern learning-based methods.
Shaofeng Zhang, Hongyuan Su, Qing Peng et al.· 0 citations
A systematic literature review on how RL are adapted and scaled as a fundamental post-training tools and how innovations in the RL pipeline enhance the domain-specific LLMs is conducted.
Qianyue Hao, Lin Chen, Xiaoqian Qi et al.· ACM Computing Surveys· 1 citation