AutoPref: Automatic Discovery of Task-Specific Preference Objectives for Neural Combinatorial Optimization
Combinatorial optimization problems (COPs) underpin many real-world decisions, but their exponentially large search spaces make high-quality solutions costly to obtain. Neural combinatorial optimization (NCO) learns fast construction policies, typically with reinforcement learning (RL), while preference-based NCO impro...