Skip to content

How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits

May 2026 · arXiv.org · Vol abs/2605.08348 · 3 citations · 48 references
Computer Science

TL;DR

Questions are raised about the degree to which circuits can support targeted understanding of, and intervention on, model behavior, by studying their consistency and specificity.

Abstract

The circuits framework in mechanistic interpretability aims to identify sparse subgraphs of model components that are causally responsible for a behavior, typically evaluated by measuring necessity and sufficiency. But these criteria say little about whether a circuit consistently captures how a model performs a task, or if it is specific to that task. We study these two properties, consistency and specificity, across six tasks and five models, extracting circuits at the component level (attention heads and MLP blocks) and at the level of individual MLP neurons. We find that component-level circuits are highly consistent and causally important on most tasks, but they are not specific: ablating one task's circuit damages another task's performance about as much as that task's own circuit does. Neuron-level circuits, on the other hand, exhibit higher task-specificity but are far less consistent within tasks. This is explained by circuit overlap: component-level circuits share most of their components across all task pairs, related or not, while neuron-level circuits overlap only between closely related tasks. In a case study of the components shared by the task circuits of Llama-3.2-3B, we show that they consist mostly of MLP blocks, while the few attention heads within turn out to be generic attention-sink heads. Overall, our findings raise questions about the degree to which circuits can support targeted understanding of, and intervention on, model behavior.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?

Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlyin...

Li Zhang, Chu-Qin Geng, Mark Zhang et al. · 0 citations
Preprint Aug 2026

Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit

Circuit Condensation is introduced, which post-trains models to concentrate behaviors into smaller causal graphs, which are smaller than the strongest frozen baseline in 30 of 32 settings, and shows that weight updates, rather than search alone, drive the reduction.

Sailesh I. S. Kumar · 0 citations
#machine learning Preprint Sep 2026

Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs

This work provides the first mechanistic base-vs-instruct comparison of conflict-resolution circuits, and believes that because the conflict circuit is preserved rather than rebuilt, interpretability and control tools calibrated on base models should transfer directly to their deployed instruct siblings.

S. Pandere, Gautam Ranka, Ritika Varshney et al. · 0 citations
#machine learning Preprint Aug 2026

How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution

Concept-Targeted Attribution (CTA) provides a framework for moving from behavioral probe accuracy to mechanistic explanations of probe performance, enabling more detailed audits of internal concept representations, including safety-critical ones.

V. Palit, Florent Draye, Terry Jingchen Zhang et al. · 0 citations
#natural language process... Preprint Aug 2026

Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning

This work systematically investigates when, where, and to what extent conditional memory should participate in scientific reasoning, and proposes a Knowledge Boundary-Aware Router that uses task-specific input proxies available before generation to determine whether memory is activated, which layer-stage nodes receive...

Zhen Bi, Xue-Shu Chen, Yan Wang et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.