Skip to content

Learning Kernels by Alignment for Multiclass Bayes Classification

Sep 2026 · 0 citations · 35 references
Computer Science

TL;DR

A principled learned-kernel framework is established that unifies representation learning, kernel alignment, and Bayesian classification, and extends naturally to the multiclass setting.

Abstract

Kernel methods separate data representation from decision-making, but typically require the kernel to be chosen in advance. We show that this kernel can instead be learned by alignment, and develop the resulting framework through the recently introduced Collaborative Learning and Inference (CLaI). We show that Collaborative Learning can be viewed as a kernel alignment process, in which an embedding is trained so that its induced similarity matches a label-derived target kernel. We also prove that Collaborative Inference is equivalent to kernel Bayes classification with Parzen-window density estimation. Motivated by these perspectives, we generalise CLaI by replacing cosine similarity with a learned Mahalanobis distance and extend it to multiclass classification. On CIFAR-10, PathMNIST, and SleepEDF, the Mahalanobis formulation improves accuracy, converges faster, and yields lower calibration error than the cosine-based variant. Auxiliary experiments further support these connections, showing that CLaI produces latent signals of the same form as a Gaussian process, while achieving competitive calibration on sepsis prediction. Together, these results establish a principled learned-kernel framework that unifies representation learning, kernel alignment, and Bayesian classification, and extends naturally to the multiclass setting.

View source

Similar papers

Preprint Sep 2026

A Weighted Kernel Method for Approximation that Adapts to Learned Multivariable Structure

This work introduces total sensitivity kernels (TSKs), a method based on families of weighted ANOVA kernels that learn and adapt to this multivariable structure, and establishes consistency of a finite-data formulation based on minimum-norm interpolation.

John Darges, Laura Weidensager · 0 citations
Preprint Aug 2026

Kernel Methods for Learning Operators with Multiple Inputs and Outputs

This work introduces a general kernel-based encoder-decoder framework for operator learning that separates observation, representation, learning, and reconstruction, and develops this framework for multi-input, multi-output operator learning, where operators map between products of potentially distinct function spaces.

Adrien Weihs, Chun-Yang Liao, Jingmin Sun et al. · 0 citations
Open access Sep 2026

Signal-Adapted Kernels and Lazy-Evaluation Gaussian Process Regression

Constant-bandwidth radial basis function (RBF) regression models are popular due to their simplicity, but they exhibit oscillatory patterns that resemble the Runge and Gibbs phenomenon on many real-world datasets. The Gaussian process regression (GPR) can be considered as an extension of RBF regression in that some R...

R. Wang, David A. Campbell · 0 citations
Preprint Aug 2026

Why not to use the Gaussian kernel

Kernels measure similarity or correlation in tasks such as regression and classification. The Gaussian kernel, other names of which include squared exponential and radial basis function kernel, is one of the most popular in Gaussian process regression. We argue that the Gaussian kernel is best avoided and should never...

Toni Karvonen, C. Oates · 0 citations
#machine learning Preprint Oct 2026

Kernel Singular Value Decomposition with Extension to Multiple Data Sources

Kernel Singular Value Decomposition (KSVD) learns a pair of singular vectors w.r.t. an asymmetric kernel matrix, which can be induced by two data sources, e.g., the queries and keys in self-attention or the rows and columns of a given matrix. In this work, we extend KSVD to multiple data sources, namely eKSVD, which co...

Xin-Jie Zeng, Qing-Hua Tao, Johan A. K. Suykens · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.