Skip to content

Learning to Optimize through Solver-Grounded Self-Play

Sep 2026 · 0 citations · 62 references
Computer Science

TL;DR

Results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.

Abstract

Optimization modeling is central to many decision-making scenarios, but traditionally requires extensive domain expertise. While Large Language Models (LLMs) have shown promise in automating this process, current training paradigms mainly rely on human-annotated or teacher-generated datasets. This dependence introduces a Generalization Ceiling, where models overfit to narrow data distributions, and Capability Anchoring, where models'reasoning is bounded by annotator proficiency and teacher model capability. In response, we propose OPT-Zero, the first fully self-play training framework for optimization modeling that requires zero external training data. OPT-Zero employs a single LLM in a dual-role closed loop: a Proposer that synthesizes increasingly challenging optimization problems alongside their mathematical formulations and solving code, and a Solver that attempts to resolve the problems given only natural-language problem descriptions. Grounded in execution feedback from external optimization solvers, we alternately train both roles using reinforcement learning. This process fosters an auto-curriculum in which the Proposer and Solver co-evolve: generating harder valid problems by the Proposer seamlessly enhances the structural reasoning ability of the Solver. Extensive results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models

SOLID is proposed, a novel framework for self-improving OR language models without verified answers or external evaluators that improves solution accuracy for both general-purpose and OR-tuned models over outcome-only group-relative training.

Rui-Chen Zhu, Ming-Long Cao, Chen-Yu Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Frontier Learning: Training LLM Reasoners at the Edge of Capability

Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal aris...

Robin Faro, S. Ramesh, Ilija Bogunovic et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Extremely Sparse Supervision Incentivizes Reasoning Ability

Large language models demonstrate increasingly strong reasoning capabilities through effective post-training. Yet, prevailing post-training methods optimize over massive numbers of tokens, implicitly assuming that effective learning must be token-intensive. We revisit this assumption in the on-policy distillation (OPD)...

Zhi-Shuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu et al. · 2 citations
Review Sep 2026

Reinforcement Learning Post-Training for Reasoning Large Language Models: Methods, Systems, and Evaluation

Reinforcement learning (RL) has become a central post-training approach for reasoning and agentic large language models (LLMs), particularly when task outcomes can be verified automatically. Comparisons across this literature remain difficult because a reported gain may combine changes to the learning signal, policy co...

Liu Yang, Han Zhu, Zheng-Yang Zhong et al. · 0 citations
#machine learning Preprint Aug 2026

Learning Generalizable Behaviors for Terminal Agents

River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization is proposed, which achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks.

Yi-Fan Yao, Bo Pang, Xuan-Phi Nguyen et al. · 1 citation
#artificial intelligence Preprint Sep 2026

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers, a cooperative framework comprising a Planning Player that generates tasks, an Execution Player that produces multi-turn trajectories with Python tool calls, and an Evaluation Player that constructs executable verifiers, highlights cooperation among specialized players as a promising path toward self-enh...

Wen-Jie Liao, Liang Zhao, Ze-Hong Cao · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.