Skip to content

The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors

Sep 2026 · 0 citations · 51 references
Computer Science Mathematics

TL;DR

The direction of ignorance is causally active: raising or lowering $\lambda$ at the final prediction state steers the prediction toward or away from the unigram prior in KL divergence, with larger models generally exhibiting lower prior reliance in the high-context limit.

Abstract

What does a language model predict when it has few clues? The answer lurks in its unembedding geometry: a single direction of the unembedding matrix encodes the unigram distribution of the training corpus, which serves as the Bayesian prior the model falls back on when uncertain. This structure --- which we term the \emph{direction of ignorance} --- appears in all four model families examined (\texttt{Llama}, \texttt{Qwen}, \texttt{Gemma}, and \texttt{Pythia}), ranging from 0.4B to 405B parameters. Projecting the final prediction state onto this direction yields a per-token \emph{prior loading factor} $\lambda$, which, empirically, declines steadily as the context becomes more informative. Formally, the same projection decomposes the prediction state into two orthogonal vectors that correspond exactly to the two factors of a tempered Bayesian update: a unigram prior raised to the exponent $\lambda$ and a context-driven likelihood. This geometric-probabilistic interpretation calibrates $\lambda$, making it meaningfully comparable across model sizes and families, with larger models generally exhibiting lower prior reliance in the high-context limit. Finally, we show that the direction of ignorance is causally active: raising or lowering $\lambda$ at the final prediction state steers the prediction toward or away from the unigram prior in KL divergence.

View source

Similar papers

Preprint Aug 2026

The Optimal Discounting Parameter of the Power Prior under Predictive Log-Loss

The power prior of Ibrahim and Chen incorporates historical data into a Bayesian analysis by raising the historical likelihood to a power $a_0 \in [0, 1]$. The choice of the exponent has remained an open question. This paper gives a closed-form answer under the predictive log-loss. For a model with $d$ parameters, a hi...

Yuriy A. Reznik · 1 citation
Preprint Aug 2026

On the relaxation problem in statistical mechanics

We reformulate the relaxation problem in statistical mechanics by making explicit what are the \emph{operational} objects subject to relaxation: the local time statistics of the recorded signal $Z(t)$. These local time statistics are simply the estimated histograms of observations $\{Z(t_i)\}_{i=1}^M$ performed at unif...

Giuseppe Del Vecchio Del Vecchio · 0 citations
Preprint Aug 2026

Attention-Path Fragility as an Uncertainty Signal in Large Language Models

It is proposed that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways, a training-free estimator that masks attention heads and measures the BALD mutual information...

Minsoo Kim, Sungyoung Ji, Kisung Moon et al. · 0 citations
#small language model Preprint Aug 2026

Planting a Latent Variable in Natural-Looking Text: a More Realistic Test of Belief States in LLMs and Their Link to Concept Geometry

This work plants a controllable latent variable inside natural-looking text and arranges the 8 states themselves on a ring, in the exact order of the Markov chain, which is supporting evidence that a concept's geometry can be formed by the statistical dynamics of the latent variable behind it.

Alexandru-Iulius Jerpelea · 0 citations
#machine learning Preprint Sep 2026

Marginal Log-Likelihood Increments under Dirichlet-Smoothed Markov Estimation

For a Dirichlet-smoothed transition model, the effect of adding one workflow trace to the training archive is an exact change in reference-weighted log likelihood. We derive that change and show that it is a weighted reduction of Kullback--Leibler divergence between the reference conditionals and the model. From this f...

Levin David Schwab · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.