Skip to content
Review Open access

Everything Is Prediction: Modern Machine Learning as Bayesian Inference

Aug 2026 · Entropy · Vol 28 · 0 citations · 108 references
Medicine

Abstract

We argue that the core methods of modern machine learning—conformal prediction, large language models and in-context learning, and generative/diffusion models—are not rivals to Bayesian inference but implementations of it, almost always of its predictive (de Finetti) form rather than its parameter-centric (prior-to-posterior) form. (1) Background: A quarter-century after Breiman contrasted the “data-modeling” and “algorithmic” cultures of statistics, we revisit that dichotomy and argue that the predictive view dissolves it—both cultures target the one-step-ahead density p(yn+1∣y1:n), differing only in how they compute it. (2) Methods: We organize the modern toolkit around this predictive object and around amortization—replacing per-dataset inference with a single map learned by simulation—using Generative Bayesian Computation (GBC) as the connective spine. (3) Results: Conformal prediction, autoregressive language models, prior-data fitted networks, and score-based diffusion are each shown to construct, calibrate, or sample from the predictive object, summarized in a single “Rosetta” table; because all are fit by proper scoring rules—equivalently, by Bregman divergences—the information-theoretic frame is the natural unifier. (4) Conclusions: The equivalence is exact in idealized limits, asymptotic under exchangeability and its martingale relaxations, and measurably approximate for trained models—a three-grade taxonomy that we make explicit, row by row, and that bounds the thesis: prediction is not attribution, and the predictive view is deliberately silent about causal structures.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.