Skip to content

BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services

Jul 2026 · arXiv.org · Vol abs/2607.11042 · 0 citations · 53 references
Computer Science

TL;DR

BackendForge is introduced, a benchmark of 56 contract-defined backend generation tasks rewritten from real open-source applications that suggests that current LLMs can implement many local API behaviors, but still struggle to produce complete backend services.

Abstract

Large language models (LLMs) are increasingly used in agentic coding settings, where they can inspect files, execute commands, run tests, observe failures, and iteratively revise code. This shift raises a central evaluation question: can an agentic LLM generate an end-to-end software artifact that is both deployable and behaviorally correct under execution? Backend services provide a controlled but realistic substrate for this evaluation. Their APIs expose application-level executable semantics, and deployed behavior can be checked deterministically against an OpenAPI contract through black-box HTTP interactions. We introduce BackendForge, a benchmark of 56 contract-defined backend generation tasks rewritten from real open-source applications. Given a visible specification and an OpenAPI contract, an LLM must generate a Dockerized service that is built, deployed, and evaluated only through HTTP tests. To strengthen evaluation without introducing hidden requirements, BackendForge uses a test agent and a code agent to co-evolve the test oracle and reference service, where the test agent proposes specification-grounded backend tests and the code agent repairs the reference implementation. Although the best-performing model, GPT-5.5, succeeds on 55.4\% of tasks under the base oracle, it succeeds on only 28.6\% under the final oracle. This gap suggests that current LLMs can implement many local API behaviors, but still struggle to produce complete backend services.

View source

Similar papers

Preprint Aug 2026

Framework and Benchmark for Code-Driven Agentic Testing in Web Development

CAT is introduced, a paradigm in which the agent writes Playwright code to drive the browser, gathers feedback, and autonomously explores web applications to uncover bugs, revealing a clear gap between current VLM capabilities and the demands of real-world testing in AI web development.

Bin Hong, Zhen-Chao Zhang, Ji-Yuan He et al. · 0 citations
Preprint Aug 2026

Kozuchi Agent: A Language-Agnostic Open-Weight Agent for Software Repair

Kozuchi Agent, a language-agnostic open-weight repair agent and CI-operated evaluation pipeline, is presented, showing that the remaining gap is primarily semantic correctness and selection rather than edit formatting or proprietary-model access.

M. Bahrami, Kosaku Kimura, Satoshi Munakata et al. · 0 citations
Preprint Aug 2026

Evaluating Agentic Code Repair Capabilities in Distributed Systems

DDBench is introduced, a code-repair benchmark of 60 historical bugs mined from 13 open-source distributed systems, partitioned into three difficulty tiers, isolating the effect of debugging context from model capability.

Yi-Bo Yan, Huijuan Wang, Jun-Zhou He et al. · 0 citations
Jul 2026

Specification-Driven DevOps for Multi-Service Environments

This study investigates whether a frontier LLM can generate Dockerfiles and Docker Compose configurations for multi-service applications using repository contents without access to developer-authored deployment artifacts and analytically derives a minimal explicit deployment specification for information that cannot be...

Oleg Grynets, Kyrylo Fursov, V. Lyashkevych et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.