Skip to content

The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Convergent Category Geometry in Small Language Models

2026 · arXiv.org · Vol abs/2607.16741 · 0 citations · 18 references
Computer Science

TL;DR

This work extends the framework that demonstrated that truth representations in large language models are universal across statement polarity but reside within a multidimensional subspace along three questions: how the dimensionality of the subspace depends on the model’s knowledge, which architectural component builds the truth direction, and what the direction is a mixture of.

View source

Similar papers

#natural language process... Preprint Aug 2026

How Language Models Organize and Structure Moral Knowledge

This work trains six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category, and examines how the resulting directions relate to each other in representation space, finding the directions neither collapse into a single moral detector nor isolate from one another.

Orion Reblitz-Richardson · 1 citation
Preprint Aug 2026

Local and Global Regimes of Geometric Complexity in Language Model Representations

A scale-dependent transition between two ID regimes is found: at low lexical diversity, conditions with fewer unique final words produce higher ID, while at high lexical diversity, this ordering reverses, and conditions with more unique words produce higher ID.

Arwa Osman, Marco Baroni, Iuri Macocco · 0 citations
#artificial intelligence Preprint Sep 2026

Rules Amortize, Pairings Don't: Linguistic Structure Determines What Latent Task Representations Can Replace In-Context Learning

In-context learning (ICL) can be amortized into latent objects (task vectors, function vectors, context vectors) that recover few-shot behavior at zero-shot inference cost, but recent theory shows a static vector acts as a single synthetic demonstration and must fail on high-rank mappings such as word-level bijections....

Gunmay Jhingran · 0 citations
#machine learning Preprint Oct 2026

Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of...

Tido Specht, Elias Krey, Nils Neukirch et al. · 0 citations
Preprint Aug 2026

Attention-Path Fragility as an Uncertainty Signal in Large Language Models

It is proposed that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways, a training-free estimator that masks attention heads and measures the BALD mutual information...

Minsoo Kim, Sungyoung Ji, Kisung Moon et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.