Skip to content
Preprint

Measuring Activation Control in Large Language Models

Aug 2026 · 0 citations · 17 references
Computer Science

TL;DR

The Activation Controllability Benchmark is introduced to quantify the extent to which models can modulate their residual stream via natural-language instruction, and suggests that control over the activation space itself could become a confound for monitoring as introspective capabilities increase.

Abstract

Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especially when evaluation-aware models exhibit scheming or deception. However, if models can also control their own activations, deception could extend into the latent space itself. With this in mind, we introduce the Activation Controllability Benchmark to quantify the extent to which models can modulate their residual stream via natural-language instruction. Across model families and capability levels, we find that most LLMs can control the direction and magnitude of their residual stream activations with some degree of temporal resolution, though performance varies considerably across models. In simple tasks, this level of control can evade activation-based monitoring methods (including linear probes, natural language autoencoders, activation oracles, and the Jacobian lens), albeit imperfectly. These results suggest that control over the activation space itself could become a confound for monitoring as introspective capabilities increase; therefore, we recommend that frontier labs and evaluators track activation controllability in future models.

View source

Similar papers

Jul 2026

Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

Activation steering controls model behavior by editing internal activations at inference time, and finds that a single MLP neuron is eval-correlated but not causal at both scales, and that scanning the real Pile yields a natural-text baseline competitive with the optimizer for the internal direction.

Deepanshu Mody, Samar Agarwal, Utkarsh Mittal et al. · 0 citations
Open access Aug 2026

Towards Trustworthy Large Language Models

An integrated conceptual frame-work that couples attention- and perturbation-based explainability with lightweight hallucination-detection signals and token-efficient inference strategies is presented, and a set of cross-cutting consistency metrics are instrumented with a set of cross-cutting consistency metrics.

Sakshi Parate, Shreyans Sanyal · 0 citations
#artificial intelligence Preprint Aug 2026

Toward Latent Language Model Skills Steering and Optimization: An Empirical Study

This empirical study investigates whether procedural LLM skills can be represented as directions in activation space and whether vector-space operations over these directions can express skill-level behaviors and finds that procedural skills admit a vector-space representation.

Xun-Yi Jiang, Junda Wu, Yuxin Xiong et al. · 0 citations
Preprint Jul 2026

Forecasting Side Effects of Activation Steering

Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, steering often produces unintended side effects on other behaviors, making it difficult to deploy safely. We therefore ask: can these side effects be forecasted before steering is applied? We answer this question by constructing a cross-effect matrix over a taxonomy of 67 behaviors across three open-weight language models. We find that side effects are common, structured, and often asymmetric, revealing interactions that cannot be explained by existing similarity-based heuristics. Despite this complexity, we show that side effects are largely predictable before steering is performed. Their magnitude depends primarily on the target behavior, while their direction can be forecasted from the model's unsteered representations with substantially higher accuracy than simple baselines. Our results demonstrate that activation steering has systematic and forecastable side effects, enabling proactive safety auditing and more informed deployment of steering interventions.

Yong-Ong Chong, Alson Wei Jie Sim, Peixin Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

Disentangling Steering Vectors

Activation steering has emerged as a lightweight, inference-time approach to control the behavior of Large Language Models (LLMs). However, traditional steering vectors used to intervene in LLMs'activations, such as those derived from the difference-in-means method, tend to entangle multiple semantic and stylistic concepts into a single composite direction, leading to unpredictable steering effects. Our core objective is to disentangle this composite direction into its constituent concepts. To this end, we propose Steering Vector Dissection, a framework to explicitly isolate individual and semantically consistent features from these composite directions. Specifically, we pair positive and negative activations and take their differences to generate a set of instance-level steering vectors, and train a dedicated Sparse Autoencoder (SAE) directly on them. Quantitative evaluations across two datasets, two models, and two intervention depths show that our method yields a set of semantically consistent basis vectors whose steering effects are mutually distinguishable. Furthermore, we show that this disentanglement enables precise control over model behaviors.

Takeru Hiramatsu, Kyohei Atarashi, Koh Takeuchi et al. · 1 citation
#natural language process... Preprint Aug 2026

Latent Mechanisms of Language Control in Multilingual Language Models

A comparative study of three methods that identify language-controlling latents in cross-layer transcoders: activation value-based selection (ValSel), activation frequency-based selection (FreqSel), and LLM-generated latent annotation-based selection (AnnSel) finds all three effectively manipulate generation language.

Ryouya Mitsuhashi, Sabri Boughorbel, Majd Hawasly · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.