Skip to content

Unsupervised Features Mining via Activation Geometry

Jul 2026 · arXiv.org · Vol abs/2607.04222 · 0 citations · 34 references
Computer Science

TL;DR

The same method is used to select the best training datasets for prompt-injection classifier probes: while similarity between ordinary activations is almost unrelated to downstream performance, RFD-based similarity achieves Top-1 and Top-2 accuracy.

Abstract

Interpretability methods aim to reveal the features represented inside large language models (LLMs). Many existing methods begin with labeled examples of a human-defined concept that may reflect human biases, and then identify how that concept is represented within the model, for example in its activation space or through other decomposition methods. We introduce \emph{Mining via Activation Geometry} (MAG), a simple unsupervised framework for extracting reasoning features from model activations by prepending the same natural-language instruction $Q$ to every input $p$, where $Q$ defines the reasoning feature of interest, such as ``Can this object be found in the desert?''or ``Is this prompt malicious?''We measure how the instruction changes the model's internal representation using $m(Q \mid p) - m(p)$ at a single readout point. We explore eight different MAGs. The extracted reasoning features predict the models'own world understanding and judgment, can be approximated into a single activation direction, we found that some features are more linearly represented and some less, this linear representation, which is vector steering, can change the LLMs'decisions through activation steering by injecting reasoning features. Finally, we use the same method to select the best training datasets for prompt-injection classifier probes: while similarity between ordinary activations is almost unrelated to downstream performance, RFD-based similarity achieves $94.7\%$ Top-1 and $100\%$ Top-2 accuracy.

View source

Similar papers

Latent Fact-Checking: Detecting Misinformation through Activation Engineering

Findings provide evidence that truthfulness is a structured, linearly separable concept in the latent space of pretrained language models, and point toward interpretability-driven misinformation detection as a practical complement to retrieval-based pipelines.

P. Barcelos, Otávio Parraga, M. M. Delucis et al. · 0 citations
#machine learning Preprint Sep 2026

ProToMEx: Rapid, Interpretable Explanations via Structured Representations

It is demonstrated empirically that ProToMEx not only produces explanations of comparable fidelity to popular methods like SHAP and LIME but also drastically reduces the amortised computational cost of generating local explanations, making it highly suitable for real-time applications.

A. Georgara, Adarsh Valoor, Sarvapali D. Ramchurn · 0 citations
Preprint Aug 2026

Deep Learning Models Also Recall Features

This paper argues that factual recall points to something broader: a general kind of operation in deep learning models, which is called feature recall, and defines it, shows it applies across architectures, and contrasts it with the established paradigm of feature combination.

P. Beckmann · 2 citations
Jul 2026

From Found to Designed: Concepts as a Design Axis for Large Language Models

This taxonomy reveals three broad patterns: inference-time approaches remain comparatively underexplored, related ideas have developed largely in isolation across pipeline stages, and externally grounded methods span the entire pipeline despite often being described under different terminology.

Sha-Ni Chen · 0 citations
Preprint Jul 2026

The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Shared Category Geometry in Small Language Models

B\"urger et al.\ (2024) demonstrated that truth representations in large language models are universal across statement polarity but reside within a multidimensional subspace. The truth value of a statement is linearly readable from a residual stream of language model, but it is not clear how much of that representatio...

Francesco Vicidomini · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.