Skip to content

Learning from 53.6K Real-World Developer Edits of AI-Generated Code

Jul 2026 · arXiv.org · Vol abs/2607.25130 · 1 citation · 59 references
Computer Science

TL;DR

DECODE (Developer Edits of Code Dataset), a dataset of 53.6K real-world in-IDE code edits of AI-generated code in Python, TypeScript, and JavaScript, is introduced, and finetuning on DECODE enables open-source 3B models to perform code edit prediction tasks significantly better than frontier LLMs.

Abstract

Imperfections in AI-generated code require that software developers modify the generated code manually, or by re-prompting an AI programming assistant. Manual code edits provide more realistic and granular information on editing behavior than Git commits, which only contain final successful code snippets. Yet, due to a lack of high-quality, realistic code editing data, LLMs are mostly trained on publicly available Git data (e.g., commits). To address this gap, we introduce DECODE (Developer Edits of Code Dataset), a dataset of 53.6K real-world in-IDE code edits of AI-generated code in Python, TypeScript, and JavaScript, sourced from 1K+ developers. First, we demonstrate the utility of DECODE for data analysis, obtaining insights on when, why, and how AI-generated code is edited. We find that most edits occur within the first 15 minutes after accepting an AI completion, resulting in the removal of AI completions in 31% of edit trajectories. Second, we use DECODE to benchmark the ability of LLMs to predict code edits. We find that finetuning on DECODE enables open-source 3B models to perform code edit prediction tasks significantly better than frontier LLMs. We then discuss implications of this work, emphasizing the necessity of developer-centric machine learning approaches for future AI programming assistants.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

What is the Difference Between Me and You? Benchmarking the Quality Gap Between Human-Written and AI-Generated Code

This work compares human-written and AI-generated code at scale, characterize structural complexity and statistical naturalness, and map static-analysis findings onto Orthogonal Defect Classification for defects and the Common Weakness Enumeration for vulnerabilities, making authors and languages directly comparable.

Cristina Improta, Pietro Liguori, D. Cotroneo · 0 citations
Conference Aug 2026

Generating Behavior-Driven Development Artifacts

Behavior-Driven Development (BDD) scenarios coexist with large volumes of semi-structured records (e.g., documentation, issues, informal feature descriptions), and keeping the two synchronized is laborintensive. We present a study of bidirectional generation 11https://github.com/Artin-Biniek/Submission-Project between...

A. Biniek, Nafiseh Kahani · 0 citations
Conference Open access 2025

The end of Coding: Predicting a Post-Code World with AI-Native Software Development

The research found out three things about how developers use Artificial Intelligence, adoption of Artificial Intelligence satisfaction, and with Artificial Intelligence the different ways developers are using Artificial Intelligence is changing.

P. Vijayakumar, Jegatheeswari Perumalsamy, P. Parida et al. · 0 citations
Open access Aug 2026

Detecting AI-Generated Text and Code: An Empirical Study of Cross-Generator and Cross-Domain Generalization

A paired-prompt benchmark for human-versus-machine detection across English text, Python code, and mixed text–code documents shows that reliable deployment requires cross-domain evaluation, mixed-content testing, and calibration beyond in-distribution accuracy.

Neethika Alluri, Pardha Saradhi Varma Gottumukkala, H. Indukuri · 0 citations
#artificial intelligence Preprint Sep 2026

Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models

Large language models used for code editing can be trained and deployed in at least two output regimes: direct generation, where the model emits the entire modified file in one shot, and iterative diff-based generation ("steps"), where the model emits a sequence of localized search/replace edits applied one at a time u...

Andrej Andrejev · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.