May 2026· arXiv.org· Vol abs/2605.08348· 3 citations· 48 references
Computer Science
TL;DR
Questions are raised about the degree to which circuits can support targeted understanding of, and intervention on, model behavior, by studying their consistency and specificity.
Abstract
The circuits framework in mechanistic interpretability aims to identify sparse subgraphs of model components that are causally responsible for a behavior, typically evaluated by measuring necessity and sufficiency. But these criteria say little about whether a circuit consistently captures how a model performs a task, or if it is specific to that task. We study these two properties, consistency and specificity, across six tasks and five models, extracting circuits at the component level (attention heads and MLP blocks) and at the level of individual MLP neurons. We find that component-level circuits are highly consistent and causally important on most tasks, but they are not specific: ablating one task's circuit damages another task's performance about as much as that task's own circuit does. Neuron-level circuits, on the other hand, exhibit higher task-specificity but are far less consistent within tasks. This is explained by circuit overlap: component-level circuits share most of their components across all task pairs, related or not, while neuron-level circuits overlap only between closely related tasks. In a case study of the components shared by the task circuits of Llama-3.2-3B, we show that they consist mostly of MLP blocks, while the few attention heads within turn out to be generic attention-sink heads. Overall, our findings raise questions about the degree to which circuits can support targeted understanding of, and intervention on, model behavior.
Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlyin...
Li Zhang, Chu-Qin Geng, Mark Zhang et al.· 0 citations
Circuit Condensation is introduced, which post-trains models to concentrate behaviors into smaller causal graphs, which are smaller than the strongest frozen baseline in 30 of 32 settings, and shows that weight updates, rather than search alone, drive the reduction.
This work provides the first mechanistic base-vs-instruct comparison of conflict-resolution circuits, and believes that because the conflict circuit is preserved rather than rebuilt, interpretability and control tools calibrated on base models should transfer directly to their deployed instruct siblings.
S. Pandere, Gautam Ranka, Ritika Varshney et al.· 0 citations
Concept-Targeted Attribution (CTA) provides a framework for moving from behavioral probe accuracy to mechanistic explanations of probe performance, enabling more detailed audits of internal concept representations, including safety-critical ones.
V. Palit, Florent Draye, Terry Jingchen Zhang et al.· 0 citations
This work systematically investigates when, where, and to what extent conditional memory should participate in scientific reasoning, and proposes a Knowledge Boundary-Aware Router that uses task-specific input proxies available before generation to determine whether memory is activated, which layer-stage nodes receive...
Zhen Bi, Xue-Shu Chen, Yan Wang et al.· 1 citation
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.