Skip to content

HyperMC: Multi-Fidelity Hyperparameter Tuning for Stochastic Gradient MCMC

Sep 2026 · 0 citations
Mathematics Computer Science

TL;DR

Experiments on logistic regression, probabilistic matrix factorization, and Bayesian neural networks show that HyperMC improves posterior approximation or predictive calibration relative to MAMBA, grid search, and heuristic baselines, while Robust HyperMC yields more stable and reproducible tuning results.

Abstract

Stochastic gradient Markov chain Monte Carlo (SGMCMC) methods enable scalable Bayesian inference, but their performance depends strongly on hyperparameters such as the step size, mini-batch size, and number of leapfrog steps. Since most SGMCMC algorithms lack a Metropolis-Hastings acceptance rate, standard acceptance-based tuning methods are not directly applicable. We propose HyperMC, a multi-fidelity tuning framework that combines Hyperband-style resource allocation with kernel Stein discrepancy (KSD) evaluation. By running multiple successive-halving brackets, HyperMC balances broad exploration of a continuous hyperparameter space with increasingly accurate evaluation of promising configurations under a fixed computational budget. We further introduce Robust HyperMC, which uses global grid initialization followed by elite-guided local refinement to reduce sensitivity to random candidate generation and noisy finite-budget evaluations. Under suitable approximation and concentration conditions for the estimated KSD, we establish that the successive-halving component selects a near-optimal configuration among the sampled candidates with high probability and derive a sufficient computational budget for successful selection. Experiments on logistic regression, probabilistic matrix factorization, and Bayesian neural networks show that HyperMC improves posterior approximation or predictive calibration relative to MAMBA, grid search, and heuristic baselines, while Robust HyperMC yields more stable and reproducible tuning results.

View source

Similar papers

Open access Sep 2026

Nii-MALA: A Fast Metropolis-Adjusted Langevin Sampler with Radial Velocity Benchmarks

Nii-MALA is introduced, a high-performance C implementation of the Metropolis-Adjusted Langevin Algorithm, which leverages Message Passing Interface for parallelization and incorporates automatic differentiation backends for gradient evaluation and benchmark it against several other codes using radial velocity data fro...

Jia-Men Jiao, Sheng Jin, Wen-Xin Jiang et al. · 0 citations
#machine learning Preprint Sep 2026

High-Probability Convergence of SGD via Batched Updates

Stochastic gradient descent (SGD) is the primary workhorse for large-scale optimization. While the average behavior of its iterates, typically characterized by mean-squared error bounds, is well-understood, obtaining high-probability guarantees for the last iterate remains challenging. Prior approaches to this problem...

Feng Zhu, Robert W. Heath, Aritra Mitra · 0 citations

FUND: Density Flow for Sampling Unnormalised Distributions

Efficient sampling from Boltzmann distributions is central to modelling complex physical systems. Markov Chain Monte Carlo (MCMC) methods suffer from critical slowing down, high autocorrelation, and poor mode-mixing, limiting their scalability. Recent advances, like Boltzmann Generators, offer a promising alternative b...

V. Kanaujia, Vipul Arora · 1 citation
#machine learning Preprint Aug 2026

Enhancing Bayesian Optimization and Active Learning Through Kernel Diversity

A unified framework, KENDO (Kernel ENsemble Disagreement-aware Operator), is proposed that integrates Ensemble Gaussian Processes (EGP) with disagreement-aware acquisition strategies and extends the approach to multi-objective optimization via random scalarization that preserves the single-optimizer conditioning struct...

Heng Zhang, Hao-Tian Xiang, Konstantinos D. Polyzos et al. · 1 citation
Preprint Aug 2026

Zeroth-Order Langevin Monte Carlo via SPSA under Noisy Function Measurements

In sampling problems, gradient-based schemes such as Langevin Monte Carlo (LMC) mix faster than non-gradient-based methods, but their applicability is limited by access to the gradient of the target log-density. In practice, gradients are often unavailable and function evaluations are noisy, e.g., stochastic simulators...

Hongbo Li, J. Spall · 0 citations
#machine learning Preprint Oct 2026

Optimal Momentum Methods for Stochastic Multilevel Compositional Optimization

This paper investigates stochastic multi-level optimization where the objective is a nested composition of several smooth non-convex functions. We assume that only stochastic estimates of the gradient and function values for each level are accessible. Consequently, obtaining an accurate estimate of the overall gradient...

Wei Jiang, Rui Yan, Si-Fan Yang et al. · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.