Skip to content

Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator

Jul 2026 · arXiv.org · Vol abs/2607.18642 · 0 citations · 43 references
Computer Science

TL;DR

Spaghetti Architect is presented, a tool that mints code datasets with the control such corpora lack and gives construct-validity evidence that the quality order moves established complexity and readability metrics, and reports baselines on a four-model open ladder.

Abstract

Mined code corpora are abundant but uncontrolled: a snippet's semantics, surface"messiness,"and difficulty are whatever the wild contained; there is no known-optimal reference to grade against; and any public sample may already sit in a model's training set. We present Spaghetti Architect, a tool that mints code datasets with the control such corpora lack. An anti-optimization transpiler maps a clean, language-agnostic JSON intermediate representation to deliberately redundant, fully-flattened programs in five languages (Python, JavaScript, Go, Java, C++); every program is compiled, run, and checked against a reference oracle, so each instance is correct by construction. The clean IR is a known-optimal reference, messiness is dialed by strictly-nested anti-pattern profiles, each instance is labelled along two orthogonal difficulty axes, intrinsic (problem size) and incidental (presentation at fixed semantics), and contamination is resisted by minting fresh variants from a private held-out seed. We give construct-validity evidence that the quality order moves established complexity and readability metrics, and report baselines on a four-model open ladder: exact match rises with scale, and the intrinsic knob collapses arithmetic-aggregation accuracy of even the strongest model to zero. Further, development-set scores equal freshly re-minted held-out counterparts within $|\Delta|\le 0.012$ (comprehension) and $\le 0.011$ (refactoring); on identical programs, refactoring equivalence ($0.73 \rightarrow 0.99$) is scale-invariant while output prediction collapses; and ablating the generator's self-annotations shows they inflate the weakest model an order of magnitude more than the strongest ($-0.173$ vs $-0.017$): the annotated ladder resolves one of three adjacent pairs where the unannotated resolves all three. Open source (MIT), dependency-free, archived under a persistent DOI.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

JevVibe: Efficient Classification-Guided Secure Code Generation

This work evaluates Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, and builds JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code...

Arshak Rezvani, Sasha Behrouzi, Ahmad Sadeghi · 0 citations
#artificial intelligence Preprint Sep 2026

Benchy: towards a universal language for task-oriented AI benchmarks

Benchy is a semantic language and execution engine for benchmarking AI programs. A benchmark is completely specified by a program, a scoring function, and a dataset, B=(P,S,D), and is separate from the AI-system taking it; a run binds the two, R=(B,AI). Benchmarks are authored as canonical YAML in which each semantic c...

Francis F. Daniel, Mauro Ibañez, Francis Perelman et al. · 0 citations
#small language model Preprint Aug 2026

A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification

A reproducible, license-aware knowledge-distillation recipe addressing the constraint of deploying a safety layer for large language models on commodity hardware by partitioning the corpus into seven safety categories aligned to a public hazard taxonomy.

Edson Rodrigues da Cruz, Paulo Ricardo Ferreira Neves, P. H. Falsetti et al. · 0 citations
#machine learning Preprint Sep 2026

Numbat: Building and Verifying a Self-Contained Machine-Learning Stack

This work reports on the construction and verification of numbat, a machine-learning stack written in one general-purpose language (Zig) with no third-party runtime dependencies, and treats a widely used reference implementation as an executable specification and verify against it at five levels.

Thang Tran, Lan Dang · 2 citations
Preprint Aug 2026

Aray: Deterministic-First Synthesis of Benign Artifacts for YARA Validation

This work presents Aray, a deterministic-first YARA interpreter and positive-fixture synthesizer, a deterministic-first YARA interpreter and positive-fixture synthesizer that validates generated fixtures against their source rules.

Emanuel C. A. Valente, Lourenço A. P. Júnior, L. Chahud et al. · 0 citations
Preprint Aug 2026

Code as Representation: A Compilable Parsing Paradigm for Academic Documents

Compilable Academic Document Parsing (CADP) is proposed, a paradigm that reconstructs a full page as contextual \LaTeX{} plus executable Python, so that structure-preserving elements and executable chart representations can be reconstructed, recompiled, and directly verified against the source page.

Rihui Jin, Jun Wang, Chen Zhu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.