Skip to content
#data science Dataset Open access

Trust, Augment, or Replace? Cross-Project Defect Attribution — Reproducibility Bundle v4.0.0: Corrected LOPO Corpus, Executed LLM Parity Baselines, and Calibration Sensitivity Analysis (1,267 Real Defects from Defects4J and BugsInPy)

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Reproducibility bundle for the manuscript "Trust, Augment, or Replace? What a Cross-Project Defect Predictor Is Worth on 1,267 Real Defects from Defects4J and BugsInPy" (Vijay Prasad Javvadi, submitted to PeerJ Computer Science, AI Application article type). The study asks whether a file-level defect predictor trained on other projects can be trusted (probability calibration), whether adding failure-time signals helps (rule and learned fusion), and whether a prompted frontier LLM given the same information replaces it (Claude Sonnet 4.6 and GPT-4o-mini under strict information parity), all under a leave-one-project-out protocol over 33 projects. Version 4.0.0 is a corrected corpus construction and supersedes all results of version 3.0.0. A code audit before resubmission found that the earlier corpus computed Defects4J process-metric features at a synthetic post-fix tag (leaking the fix into the features), used proxy features derived from fix-pool frequencies rather than git history for BugsInPy, and applied two different distractor schemes. In this version every feature is computed from git history at the original pre-fix revision for both languages, distractors are same-directory sibling source files at that revision for both corpora, and events with no sibling are skipped. The corpus has 1,267 events (774 Defects4J bugs in 16 projects, 493 BugsInPy bugs in 17) and 6,716 candidate-file rows. All experiments were rerun: the cross-project predictor ranks a fix file first in 54% of events (ECE 0.025), the rule fusion ranks identically by construction and triples the calibration error, the learned fusion loses 4.5 points of Precision@1, and Claude Sonnet 4.6 matches the predictor pooled only through Lang, Math and Time, trailing by 7.6 points on the other 30 projects. The 3.0.0 record remains available for comparison but its numbers must not be cited. Contents: the derived event table (real_events_v4.parquet) and its build log; the corpus-construction script, the leave-one-project-out harness, the LLM baseline runner, the distractor sensitivity script, and the analysis and figure scripts that regenerate every manuscript number and figure; per-row scores, per-event metrics, summaries, the sensitivity table, and the complete per-event LLM run logs (prompt and raw response, 1,267 per model); the four manuscript figures; the manuscript draft; README in the PeerJ layout, CHANGELOG and CITATION.cff. Defects4J v2.0.1 (github.com/rjust/defects4j) and BugsInPy (github.com/soarsmu/BugsInPy) are not redistributed; the scripts rebuild the corpus from them. Code is MIT-licensed; derived data, results and figures are CC BY 4.0. GitHub mirror: github.com/javvadivijayprasad/DefectAnalysisResearch, tag v4.0.0.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Trajectory Balance: Improved Credit Assignment in GFlowNets

It is proved that any global minimizer of the trajectory balance objective can define a policy that samples exactly from the target distribution, and empirically demonstrate the benefits of the trajectories balance objective for GFlowNet convergence, diversity of generated samples, and robustness to long action sequenc...

Esmeralda S. Whitammer, Moksh Jain, Emmanuel Bengio et al. · 302 citations · ⚡60

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.