The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors
The direction of ignorance is causally active: raising or lowering $\lambda$ at the final prediction state steers the prediction toward or away from the unigram prior in KL divergence, with larger models generally exhibiting lower prior reliance in the high-context limit.