Skip to content
Preprint

Approximating Bayesian leave-one-group-out cross-validation

Sep 2026 · 0 citations
Mathematics

Abstract

When data are grouped, hierarchical or multilevel models are commonly used to account for group-level variation with group-specific parameters. Leave-one-group-out cross-validation (LOGO-CV) is a suitable tool for evaluating predictive performance for new groups, providing an estimator of the expected log predictive density (elpd). Brute-force LOGO-CV requires one model refit per held-out group, often using computationally expensive inference algorithms such as MCMC. This is costly, particularly for large numbers of groups or complex model structures. Commonly used importance sampling approximations, intended to reduce this cost, tend to fail because the group-specific parameters of the held-out group must be integrated out. We identify two key challenges in LOGO-CV elpd estimation: approximating the LOGO posterior and computing the grouped marginal likelihood. We compare 11 strategies, including 5 newly proposed, to address them. Among others, we combine Pareto-smoothed importance sampling or adaptive importance sampling with integration techniques such as Laplace approximation, adaptive Gauss-Hermite quadrature, and bridge sampling. We evaluate these strategies in both simulation experiments and real-world case studies, which show that marginalising over the group-specific parameters substantially improves the reliability of the importance sampling approaches.

View source

Similar papers

Preprint Sep 2026

Deterministic Leave-One-Cluster-Out Cross-Validation for Multilevel Bayesian Structural Equation Models

We introduce a closed-form, refit-free procedure for leave-one-cluster-out (LOCO) cross-validation in multilevel Gaussian Bayesian structural equation models (SEMs), together with predictive scoring of every nested submodel. Conditional independence of clusters given the parameters expresses the cluster-deleted posteri...

Mohammad Alhyari, Haziq Jamil, Hans Montcho et al. · 0 citations
Preprint Sep 2026

Cross Validation for the log Gaussian Cox Process

The log Gaussian Cox Process (LGCP) is one of the most widely used models for the analysis of spatial point patterns. Although Bayesian methods and software for fitting LGCPs are now well established, practical tools for model criticism, predictive assessment, and model comparison remain comparatively underdeveloped. T...

Hans Montcho, Håvard Rue, F. Lindgren et al. · 0 citations
#machine learning Preprint Sep 2026

Prior-Amortized In-Context Bayesian Inference for Generalized Linear Mixed-Effects Models

Hierarchical data is ubiquitous in the empirical sciences and is most commonly analyzed with generalized linear mixed-effects models (GLMMs). Bayesian inference for GLMMs yields calibrated uncertainty but requires MCMC; the No-U-Turn Sampler (NUTS) is the gold standard but is slow and must restart from scratch for ever...

Alex Kipnis, Marcel Binz, Eric Schulz · 0 citations
Preprint Aug 2026

Goodness-of-Fit Tests and Calibration Machine-Learning Algorithms for Logistic Regression with Sparse Data

Assessing the goodness-of-fit of a logistic regression model is a critical prerequisite before the model is used for inference. However, goodness-of-fit (GOF) tests such as the chi-square and deviance tests often give invalid results when the data are"sparse"-- a common issue with continuous predictors like age or weig...

Ebrahim Khaled Ebrahim · 1 citation
Preprint Sep 2026

Bagged Martingale Posteriors: Calibrated Uncertainty Quantification for Predictive Resampling

Martingale posteriors and related predictive resampling methods replace the likelihood--prior pair used within Bayesian inference with a predictive model for future observations. These methods are simple to implement and increasingly popular due to their computational efficiency, but little is known about their ability...

Hui Wang, Edwin Fong, David T. Frazier · 0 citations
Preprint Sep 2026

Covariate-localized False Discovery Rates

Several procedures for estimating and thresholding the local false discovery rate are introduced, and it is shown that this holds for fixed and randomized hypothesis labels, indicating that the proposed methods perform well under both frequentist and Bayesian interpretations of multiple testing.

Jonathan Lin, S. Tokdar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.