DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories
DeltaML-Bench is introduced, a benchmark comprising 48 tasks sourced from research papers that require agents to improve published baselines within imperfect, open-source repositories that indicates that scaffolding design and integrity checks are important considerations when deploying agents for autonomous ML experim...