Trust, Augment, or Replace? Cross-Project Defect Attribution — Reproducibility Bundle v4.0.0: Corrected LOPO Corpus, Executed LLM Parity Baselines, and Calibration Sensitivity Analysis (1,267 Real Defects from Defects4J and BugsInPy)
Abstract
Reproducibility bundle for the manuscript "Trust, Augment, or Replace? What a Cross-Project Defect Predictor Is Worth on 1,267 Real Defects from Defects4J and BugsInPy" (Vijay Prasad Javvadi, submitted to PeerJ Computer Science, AI Application article type). The study asks whether a file-level defect predictor trained on other projects can be trusted (probability calibration), whether adding failure-time signals helps (rule and learned fusion), and whether a prompted frontier LLM given the same information replaces it (Claude Sonnet 4.6 and GPT-4o-mini under strict information parity), all under a leave-one-project-out protocol over 33 projects. Version 4.0.0 is a corrected corpus construction and supersedes all results of version 3.0.0. A code audit before resubmission found that the earlier corpus computed Defects4J process-metric features at a synthetic post-fix tag (leaking the fix into the features), used proxy features derived from fix-pool frequencies rather than git history for BugsInPy, and applied two different distractor schemes. In this version every feature is computed from git history at the original pre-fix revision for both languages, distractors are same-directory sibling source files at that revision for both corpora, and events with no sibling are skipped. The corpus has 1,267 events (774 Defects4J bugs in 16 projects, 493 BugsInPy bugs in 17) and 6,716 candidate-file rows. All experiments were rerun: the cross-project predictor ranks a fix file first in 54% of events (ECE 0.025), the rule fusion ranks identically by construction and triples the calibration error, the learned fusion loses 4.5 points of Precision@1, and Claude Sonnet 4.6 matches the predictor pooled only through Lang, Math and Time, trailing by 7.6 points on the other 30 projects. The 3.0.0 record remains available for comparison but its numbers must not be cited. Contents: the derived event table (real_events_v4.parquet) and its build log; the corpus-construction script, the leave-one-project-out harness, the LLM baseline runner, the distractor sensitivity script, and the analysis and figure scripts that regenerate every manuscript number and figure; per-row scores, per-event metrics, summaries, the sensitivity table, and the complete per-event LLM run logs (prompt and raw response, 1,267 per model); the four manuscript figures; the manuscript draft; README in the PeerJ layout, CHANGELOG and CITATION.cff. Defects4J v2.0.1 (github.com/rjust/defects4j) and BugsInPy (github.com/soarsmu/BugsInPy) are not redistributed; the scripts rebuild the corpus from them. Code is MIT-licensed; derived data, results and figures are CC BY 4.0. GitHub mirror: github.com/javvadivijayprasad/DefectAnalysisResearch, tag v4.0.0.