Skip to content
Preprint

MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning

Aug 2026 · 0 citations · 37 references
Computer Science Mathematics

TL;DR

This paper proposes MOON (Multi-Objective OrthoNormalized Updates), which performs gradient manipulation under spectral--nuclear norm geometry and uses the orthonormalized manipulated gradient for parameter updates and empirical results show that MOON consistently improves both optimization efficiency and final multi-task performance.

Abstract

Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradient manipulation. However, most existing methods flatten model parameters into vectors and perform gradient manipulation under Euclidean geometry, thereby overlooking the matrix structure prevalent in modern architectures such as Transformers. In this paper, we show that gradient manipulation in Euclidean space does not generally yield the steepest descent direction under matrix geometry, potentially limiting optimization efficiency. Drawing from the theory of steepest descent for matrix-valued parameters, we propose MOON (Multi-Objective OrthoNormalized Updates), which performs gradient manipulation under spectral--nuclear norm geometry and uses the orthonormalized manipulated gradient for parameter updates. Theoretically, for smooth non-convex objectives, we establish convergence of the averaged Pareto-stationarity measure at rates of $\mathcal{O}(T^{-1/2})$ in the deterministic setting and $\mathcal{O}(T^{-1/4})$ under stochastic gradients. Empirical results across various benchmarks show that MOON consistently improves both optimization efficiency and final multi-task performance. Our code is available at https://github.com/KunlinLyu/MOON.

View source

Similar papers

#artificial intelligence Book Open access Aug 2026

SIMS: Scale-Invariant Merit-Function-Based Scalarization for Multi-Task Learning

SIMS adopts a transformation-induced merit function to convert the MOO problem of MTL to a single objective that renders optimization invariant to the magnitudes of losses, and proves that the requirement for scale invariance uniquely determines this transformation to be logarithmic.

Ze-Bin Chen, Fei Xing, Yang Chen et al. · 0 citations
Preprint Aug 2026

RODE: A Radial-Orthogonal Decoupled Engine for Optimization

This work introduces RODE, which gives the radial and directional components separate update rules and step sizes in the matrix Frobenius norm, and suggests that decoupling radial and directional dynamics offers a more effective and controllable approach to matrix optimization.

Guo-Xiang Xu, Bince Qu, Qi Sun et al. · 0 citations
#machine learning Preprint Sep 2026

Nonsmooth Optimization via Orthogonalized Momentum

Modern real application problems involve matrix-valued parameters, yet conventional optimizers treat them as vectors, thereby motivating matrix-aware methods that exploit input-output geometry, such as Muon which orthogonalizes the momentum matrices before parameter updates. Its empirical success raises a conceptual qu...

Lexiao Lai, Tianyi Lin, Jia-Yu Zhang · 0 citations
#machine learning Preprint Sep 2026

No-Regret Bayesian Optimization with Finite-Library Input-Warped Kernels

Gaussian-process Bayesian optimization (GP-BO) excels at black-box optimization of costly functions, e.g., hyperparameter optimization (HPO) and multi-agent system (MAS) design. Convergence-rate guarantees exist for select methods, notably GP upper confidence bound (GP-UCB), but require a fixed kernel. Critically, the...

Edvin Ketabati Augustinsson, Robert A. Bridges · 0 citations
Preprint Aug 2026

Parallelizable Gradient-Based Optimization For Multi-Objective MaxCut

This paper develops a differentiable framework for multi-objective MaxCut by combining an adjacency-based quadratic formulation with linear scalarization, thereby reducing the problem to a preference-conditioned single-objective signed-weight MaxCut problem.

Jing-Hang Huang, Alvaro Velasquez, Jia Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Better Spur Should Start From Each Objective

This work proposes Multi-Marginal Preference Optimization (MMPO), a fine-grained framework that intervenes at the data, gradient, and constraint levels rather than relying on coarse-grained global scalarization to address optimization conflicts among multiple objectives in real-world deployment scenarios.

Shang-Wen Mao, Hao Zhang, Guangtao Nie et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.