Skip to content
Open access

Deployment-Oriented Evaluation of LLM-Generated Optimization Code: Repairability, Compatibility, and Constraint-Coverage in Parallel-Machine Scheduling

2026 · IEEE Access · Vol 14, pp. 138842-138864 · 0 citations · 22 references

Abstract

Large language models (LLMs) increasingly generate executable optimization code, yet evaluations often rank programs by objective value, overlooking deployment-relevant properties such as validity under structural constraint changes, failure localization, repairability, and component compatibility. We introduce GenSE-Scheduler, a deployment-oriented framework for evaluating LLM-generated parallel-machine scheduling code. It compares monolithic schedulers with four-stage modular pipelines under typed contracts, sandboxed execution, a failure taxonomy, stage-swap repair, and pre-composition compatibility prediction. We evaluate 1,350 artifacts from four locally served open-weight model families across 1,039,600 artifact–instance pairs. Modular pipelines exhibit substantially greater outcome-level behavioral diversity than monolithic artifacts (Cliff’s <inline-formula> <tex-math notation="LaTeX">$\delta = +0.799$ </tex-math></inline-formula> for cluster entropy; <inline-formula> <tex-math notation="LaTeX">$\delta = +1.000$ </tex-math></inline-formula> for pairwise regret distance). However, modular generation does not meet the pre-registered regret-equivalence criterion (<inline-formula> <tex-math notation="LaTeX">$\Delta = +0.565 \gt \varepsilon = 0.05$ </tex-math></inline-formula>), although the effect is negligible (<inline-formula> <tex-math notation="LaTeX">$\delta = +0.044$ </tex-math></inline-formula>) among passing in-distribution pairs. Among failing modular pipelines, 97 of 98 (99.0%) admit at least one valid single-stage substitution, but this repairability is contract-level: the resulting swaps rarely meet tight scheduling-quality tolerances. Repairability is also strongly stage-dependent, with the <monospace>assign</monospace> stage as the main bottleneck. A static Module Compatibility Predictor reaches macro-F<inline-formula> <tex-math notation="LaTeX">$1=0.872$ </tex-math></inline-formula> for predicting compatible swaps before composition. Machine-eligibility stress tests—using a constraint that is present in the generation prompts but never instantiated in the C1–C4 calibration regime—yield a 0% pass rate in both an isolated eligibility class and a combined release-date/setup-time/eligibility hold-out, which we interpret as a joint calibration-coverage and serialization-contract boundary, not evidence that LLMs cannot implement eligibility constraints when those constraints are explicitly instantiated and exercised. Against a classical greedy min-completion-time list-scheduling baseline, neither architecture is makespan-competitive: the heuristic attains far lower regret and is matched or beaten by well under 1% of passing artifact–instance pairs. Overall, modular generation does not dominate monolithic generation in raw scheduling quality; rather, modular decomposition makes deployment-relevant properties measurable.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.