This work extends the framework that demonstrated that truth representations in large language models are universal across statement polarity but reside within a multidimensional subspace along three questions: how the dimensionality of the subspace depends on the model’s knowledge, which architectural component builds the truth direction, and what the direction is a mixture of.
This work trains six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category, and examines how the resulting directions relate to each other in representation space, finding the directions neither collapse into a single moral detector nor isolate from one another.
A scale-dependent transition between two ID regimes is found: at low lexical diversity, conditions with fewer unique final words produce higher ID, while at high lexical diversity, this ordering reverses, and conditions with more unique words produce higher ID.
Arwa Osman, Marco Baroni, Iuri Macocco· 0 citations
Although performance of language-based models is improved by scaling, whether the gap to a structure-aware architecture can eventually be eliminated remains untested.
In-context learning (ICL) can be amortized into latent objects (task vectors, function vectors, context vectors) that recover few-shot behavior at zero-shot inference cost, but recent theory shows a static vector acts as a single synthetic demonstration and must fail on high-rank mappings such as word-level bijections....
Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of...
Tido Specht, Elias Krey, Nils Neukirch et al.· 0 citations
It is proposed that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways, a training-free estimator that masks attention heads and measures the BALD mutual information...
Minsoo Kim, Sungyoung Ji, Kisung Moon et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.