Skip to content

What does a Bayes-filtered transformer believe? A predictive Monte Carlo approach

Jul 2026 · arXiv.org · Vol abs/2607.17060 · 2 citations · 39 references
Computer Science

TL;DR

P predictive Monte Carlo (PMC) is used as a general interpretability tool for any BFT: using only next-token generation, PMC returns an approximation to the implicit prior and posterior over the latent task, answering the interpretive question directly in latent space.

Abstract

A Bayes-filtered transformer (BFT) is a transformer trained on sequences that are generated in two steps: first a latent task is drawn from a prior, then observations are drawn conditional on that task. Trained under autoregressive log loss, the BFT's next-token prediction, in the idealized limit, is the Bayesian posterior predictive distribution (PPD) induced by that prior and that conditional law. In practice the trained BFT is only an approximation of this ideal PPD, raising an interpretive question: what prior and posterior over the latent task has the trained BFT actually internalized? Existing work answers this question by comparing the trained BFT's predictions against the predictions of various"reference"posteriors, each standing in for a different candidate algorithm or computation the BFT might be implementing. This prediction-space comparison is fragile: different posteriors can share the same posterior-mean predictions. We use predictive Monte Carlo (PMC) as a general interpretability tool for any BFT: using only next-token generation, PMC returns an approximation to the implicit prior and posterior over the latent task, answering the interpretive question directly in latent space. We apply PMC to three stylized task families spanning 0-Markov and 1-Markov exchangeability. The phenomena previously reported in these settings remain visible in latent space. Code is available at https://github.com/afiq-aswadi/bft-pmc

View source

Similar papers

Preprint Aug 2026

Stochastic Bayes factors: why, when, and how

The Bayes factor (BF) is a central tool in Bayesian hypothesis testing and model selection, yet its practical use is often challenged. Classical BFs depend heavily on prior specification, cannot be applied with improper priors, and are typically interpreted through arbitrary evidence scales. Moreover, they fail to capt...

L. Egidi, I. Ntzoufras · 0 citations
Jul 2026

BayesAME: Bayesian Active Model Evaluation

This comprehensive analysis addresses recent skepticism in the literature, establishing that non-random coreset selection is advantageous over random selection and highlighting that leveraging continuous response log-likelihoods over traditional binary scores significantly enhances estimation accuracy.

Paula Cordero Encinar, taylan. cemgil, Arnaud Doucet et al. · 0 citations
Review Open access Aug 2026

Everything Is Prediction: Modern Machine Learning as Bayesian Inference

We argue that the core methods of modern machine learning—conformal prediction, large language models and in-context learning, and generative/diffusion models—are not rivals to Bayesian inference but implementations of it, almost always of its predictive (de Finetti) form rather than its parameter-centric (prior-to-pos...

Nicholas G. Polson, Vadim Sokolov, R. Soyer · 0 citations
Review Aug 2026

Bayesian Inference Procedures for A/B Testing: An Overview

Simulations against group-sequential and always-valid frequentist baselines show that flat-prior posterior stopping exactly reproduces naive peeking, that a well-calibrated empirical Bayes prior achieves the lowest estimation error, and that expected-loss stopping minimizes regret only when shipping a null-effect varia...

M. Schultzberg, Mattias Frånberg · 0 citations
Preprint Sep 2026

Empirical Bayes for compound adaptive experiments

We investigate Empirical Bayes (EB) methods in the context of compound adaptive experiments, where the arm distribution in each experiment follows a normal distribution with an unknown mean that we seek to estimate. There are two main EB strategies: $g$-modeling, which estimates the prior by maximizing the marginal lik...

Karun Adusumilli, Jia-Ying Gu, Jun-Fan Tao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.