Skip to content

PAC-Bayes Bounds on Quotient Parameter Spaces: Geometry-induced Implicit-Bias Priors

Jul 2026 · arXiv.org · Vol abs/2607.18422 · 0 citations · 16 references
Computer Science Mathematics

TL;DR

It is shown that PAC--Bayesian analysis should be performed on the quotient predictor space: pushing a prior and posterior to the quotient preserves the empirical and population Gibbs risks while removing the nonnegative KL contribution caused solely by how the two distributions differ among parameterizations of the same predictor.

Abstract

Overparameterized models often have continuous parameter symmetries, so different parameters define the same predictor. We show that PAC--Bayesian analysis should be performed on the quotient predictor space: pushing a prior and posterior to the quotient preserves the empirical and population Gibbs risks while removing the nonnegative KL contribution caused solely by how the two distributions differ among parameterizations of the same predictor. Quotienting alone does not determine which prior to use. We construct a canonical choice of one parameterization for each predictor and account for the geometric volume of its equivalent parameterizations. This transforms a neutral reference prior into a data-independent prior that reflects the model's implicit bias. It approximates the ideal but inadmissible posterior-matched prior, which would minimize the KL term by depending on the training data. The resulting certificate is tighter exactly when this geometry-induced prior has smaller KL divergence from the learned quotient posterior than the neutral prior. We test this prediction in Fourier regression with a Hadamard parameterization and in Query-Key attention, using ordinary SGD without an explicit regularizer. The implicit-bias prior reduces the mean quotient-space KL by \(40.69\%\) and the mean PAC--Bayes certificate by \(21.40\%\) in the Fourier-Hadamard experiment. The smaller, prior-scale-dependent improvement in Query-Key attention confirms the predicted conditional nature of the effect.

View source

Similar papers

Preprint Sep 2026

On Prior-to-Posterior Stability in the Wasserstein Metric for Bayesian Inverse Problems

Priors in Bayesian inverse problems are often approximated through discretization, hyperparameter estimation, or generative modeling. Understanding how prior approximation errors propagate to the posterior and subsequent predictions is therefore important. In this work, we study the stability of the prior-to-posterior...

Liang-Hao Cao · 0 citations
Preprint Aug 2026

PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition

PAC-Bayes theory provides generalization guarantees by controlling the Kullback--Leibler (KL) divergence between posterior and prior distributions over a chosen hypothesis representation. However, predictive risk depends only on the predictive behavior induced by a hypothesis, not on the particular internal realization...

Vasant Honavar, Satish Kumar Keshri, Neil Ashtekar et al. · 0 citations
Preprint Aug 2026

Duality and Error for Predictively Oriented Inference

This work derives a finite-dimensional dual formulation of PrO inference that separates sampling fluctuation, approximation under a divergence budget, regularization, and numerical optimization error and uses an exactly solvable categorical example to show that predictive-risk convergence can imply convergence to a uni...

Aurya Javeed, D. Kouri, Teresa Portone et al. · 0 citations
Preprint Aug 2026

Exact Algebraic Computation of Learning Coefficients for Two-Dimensional Singular Models

This work presents the first deterministic algorithm that computes local RLCTs exactly for any two-dimensional model whose Kullback-Leibler distance is contact equivalent to a polynomial, derive a bound on its complexity, and demonstrate its effectiveness for a broad class of models, with applications including polynom...

Grégoire Sergeant-Perthuis, E. Tsigaridas, Jules Tsukahara Cqsb et al. · 0 citations
Preprint Aug 2026

Posterior contraction rates in Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families

We study posterior contraction in positive-order Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families. We embed the natural parameter in a Hilbert scale and model it via a standard Gaussian series prior expanded in the eigenbasis generating the scale. Under a two-sided link con...

Emanuele Dolera, Stefano Favaro, M. Giordano · 0 citations
#machine learning Preprint Sep 2026

Generalized Score Matching for Parameter Estimation on Convex Domains

This work derives the generalized score matching objective on a convex subset of $\mathbb{R}^{d}$ constructively starting from Minimum Probability Flow (MPF) learning, and shows how classical score matching as well as domain-adapted variants for non-negative data arise naturally within the proposed framework.

Nishanth Shetty, Saisuchith Mahajan, C. Seelamantula · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.