Skip to content
Preprint

DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories

Aug 2026 · 2 citations · ⚡ 1 influential · 17 references
Computer Science

TL;DR

DeltaML-Bench is introduced, a benchmark comprising 48 tasks sourced from research papers that require agents to improve published baselines within imperfect, open-source repositories that indicates that scaffolding design and integrity checks are important considerations when deploying agents for autonomous ML experimentation.

Abstract

Autonomous agents for machine learning experimentation must navigate heterogeneous repositories, repair training pipelines, and evaluate candidate improvements under realistic compute constraints. Existing benchmarks only partially capture these conditions. We introduce DeltaML-Bench, a benchmark comprising 48 tasks sourced from research papers that require agents to improve published baselines within imperfect, open-source repositories. We evaluate GPT-5 and Claude Sonnet 4 with a standard Modular agent and a search-based ARG scaffolding. In the 4 x 6h allocation, ARG raises GPT-5's per-run success rate from 9.4% to 33.9%; in the 2 x 12h allocation, GPT-5 ARG reaches 49.0%. Modular configurations exhibit specification gaming rates as high as 47.9%, while no gaming is observed in the evaluated ARG configurations. These results indicate that scaffolding design and integrity checks are important considerations when deploying agents for autonomous ML experimentation.

View source

Similar papers

Jul 2026

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

AgentHPOBench, a sequential benchmark comprising 30 executable machine learning tasks across seven research categories, shows that current agents exhibit measurable experimental optimization ability across domains, but still face clear limitations in sustained iterative refinement, complex log diagnosis, and consistent...

Tianyu Huai, Tingshuo Fan, Xinchi Chen et al. · 0 citations
#artificial intelligence Preprint Aug 2026

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

This work introduces TraceML, which pairs human and agent work on the same competitions under one version-level schema, and releases the corpus, the schema, the labelers, and the extraction pipeline at https://huggingface.co/datasets/jerryyan/TraceML.

J. Yan, Weiwei Sun, Si-Jie Li et al. · 0 citations

Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

Reproducibility is essential to scientific progress, yet the growing volume and complexity of scientific publications make exhaustive manual verification increasingly impractical. Although recent advances in large language model (LLM) agents enable automated experiment reproduction, existing evaluations largely focus o...

Han-Hua Hong, Yi-Zhi Li, H. Luu et al. · 1 citation
Preprint Aug 2026

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

The proposed AutoSaddler, an automatic harness optimization framework that formulates harness improvement as an offline learning problem and iteratively updates the harness using failure signals from mini-batches, suggests that automatic harness optimization is a promising path toward more performant and reliable agent...

Sungho Park, Wonjoong Kim, Rongyuan Tan et al. · 14 citations
#artificial intelligence Preprint Aug 2026

K-Bench: measuring model performance on real scientific agent requests

K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs is reported, arguing that the informative quantity for scientific agents is not a leaderboard position but the joi...

Aubrey M. Brueckner, Darshil Patel, Yu-Huan He et al. · 2 citations
Preprint Aug 2026

Evo-Bench: Can Language Models Improve Agent Harness?

Evo-Bench is the first benchmark designed to evaluate models'intrinsic harness-evolving capabilities across Search, Office, and General agent domains, and exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consis...

Lisheng Huang, Chen Yang, Hao Zhou et al. · 6 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.