Unembedding Geometry Reveals How LLMs Use Bayesian Priors
September 4, 2026
Analysis of Llama, Qwen, Gemma, and Pythia models shows that a specific direction in the unembedding matrix encodes the unigram training distribution. This 'direction of ignorance' allows models to perform tempered Bayesian updates, transitioning from unigram priors to context-driven likelihoods.
HOW THIS AFFECTS YOU
●
researcherYou can quantify a model's uncertainty by projecting prediction states onto the identified direction of ignorance.