Deployment-Oriented Evaluation of LLM-Generated Optimization Code: Repairability, Compatibility, and Constraint-Coverage in Parallel-Machine Scheduling
Large language models (LLMs) increasingly generate executable optimization code, yet evaluations often rank programs by objective value, overlooking deployment-relevant properties such as validity under structural constraint changes, failure localization, repairability, and component compatibility. We introduce GenSE-S...