Skip to content

Verbalized Particle Posterior: Bayesian Inference over Natural Language Hypotheses

Jul 2026 · arXiv.org · Vol abs/2607.22961 · 0 citations · 41 references
Computer Science

TL;DR

The Verbalized Particle Posterior (VPP) is proposed, which treats verbalized learning as a Bayesian inference problem: maintain a population of natural-language hypotheses as particles, update them with Metropolis-Hastings or Sequential Monte Carlo, and predict by Bayesian model averaging.

Abstract

Verbalized Machine Learning (VML) parameterizes a model as a natural-language prompt that an LLM evaluates as f(x; theta). The framework is interpretable, but it commits to a single hypothesis with no measure of uncertainty, and that hypothesis varies substantially across optimization runs on the same data. We propose the Verbalized Particle Posterior (VPP), which treats verbalized learning as a Bayesian inference problem: maintain a population of natural-language hypotheses as particles, update them with Metropolis-Hastings (VPP-MH) or Sequential Monte Carlo (VPP-SMC), and predict by Bayesian model averaging. Both algorithms treat the LLM as a black box, requiring no access to logits or gradients. A distinctive consequence follows. In classical Bayesian learning, model selection sits outside the posterior; in VPP both model structure and parameters share a single language space, and the posterior ranges over both. We evaluate VPP on regression, classification, and rule-discovery benchmarks. It improves over a single VML run on every benchmark and matches or exceeds an oracle-best ensemble of independent VML runs on most, while eliminating the catastrophic single-run failures that VML occasionally produces. Because each particle is a human-readable hypothesis, the posterior is itself something a reader can inspect, seeing in plain text which explanations the data supported and which it ruled out.

View source

Similar papers

Preprint Aug 2026

Comment on"Modeling rapid language learning by distilling Bayesian priors into artificial neural networks"

McCoy&Griffiths (2025, henceforth M&G) suggest that a Bayesian prior can be distilled into Artificial Neural Networks (ANNs) through Model-Agnostic Meta-Learning (MAML, Finn et al., 2017). They support this empirically by showing that meta-trained networks demonstrate formal language learning abilities comparable to Ya...

Orr Well, Idan Tarshish, N. Lan et al. · 0 citations
#machine learning Preprint Sep 2026

Prior-Amortized In-Context Bayesian Inference for Generalized Linear Mixed-Effects Models

Hierarchical data is ubiquitous in the empirical sciences and is most commonly analyzed with generalized linear mixed-effects models (GLMMs). Bayesian inference for GLMMs yields calibrated uncertainty but requires MCMC; the No-U-Turn Sampler (NUTS) is the gold standard but is slow and must restart from scratch for ever...

Alex Kipnis, Marcel Binz, Eric Schulz · 0 citations
#machine learning Preprint Sep 2026

ABSOL: Aggregated Bayesian Subsampling Orchestrated with LLMs

A hybrid LLM-guided Bayesian network structure-learning framework that uses LLMs as bounded semantic guides, and shows that language-derived semantic knowledge can substantially improve scalable probabilistic structure learning when used as bounded guidance within a statistically grounded reasoning pipeline.

Jackson Hassell, Chen Shen, Estevam Hruschka · 0 citations
#artificial intelligence Preprint Sep 2026

The Rise of Verbal Reinforcement Learning

This taxonomy shows how verbal reinforcement is reshaping agent development, while also defining the challenges and opportunities for building more capable and aligned agents.

Kshitij Tayal, Arun Sharma, Genta Indra Winata et al. · 0 citations
#machine learning Preprint Sep 2026

Large Language Models Develop Belief State Geometry In-Context

Large language models trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood, and representation-level evidence that ICL in open-source LLMs approximates optimal Bayesian prediction over a context-inferred generative mod...

Daniel Balcells, Andrew Lee, Chirag Rastogi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors

The direction of ignorance is causally active: raising or lowering $\lambda$ at the final prediction state steers the prediction toward or away from the unigram prior in KL divergence, with larger models generally exhibiting lower prior reliance in the high-context limit.

To-Ni J. B. Liu, Jiajun Bao, Yizhou Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.