A Better Spur Should Start From Each Objective
This work proposes Multi-Marginal Preference Optimization (MMPO), a fine-grained framework that intervenes at the data, gradient, and constraint levels rather than relying on coarse-grained global scalarization to address optimization conflicts among multiple objectives in real-world deployment scenarios.