Skip to content

Sparse Inter-Layer Dependencies of Transformer FFN Neurons

Jul 2026 · arXiv.org · Vol abs/2607.11990 · 0 citations · 38 references
Computer Science

TL;DR

A training-free attribution method that estimates the relative influence of upstream neurons and attention outputs on a target neuron's activation and identifies candidate sparse pathways with potential implications for efficient inference is introduced.

Abstract

Feedforward network (FFN) blocks account for a large fraction of the parameters and computation in Transformer architectures, yet their internal structure remains difficult to interpret due to the additive superposition induced by the residual stream. We examine whether the activation of an FFN neuron can be explained by a sparse set of preceding neuron activations and attention outputs. We introduce a training-free attribution method that estimates the relative influence of upstream neurons and attention outputs on a target neuron's activation. Empirically, across models and layers, we find that small subsets of preceding activations and attention outputs suffice to preserve neuron activations with high fidelity when all remaining inputs are masked with their average values. Effective sparsity is even greater when accounting for the inherent activation sparsity of upstream layers. Moreover, applying the neuron-specific masks in all layers simultaneously, such that the induced deviations propagate through the network, leaves model perplexity largely unchanged at moderate sparsity levels. These results demonstrate that, despite dense parameterization, FFNs exhibit sparse and structured inter-layer dependencies at the neuron level. Our method provides a practical, scalable tool for circuit-level interpretability and identifies candidate sparse pathways with potential implications for efficient inference.

View source

Similar papers

#small language model Preprint Aug 2026

The Von-Neumann State-Space Transformer for neural decoding

A von-Neumann inspired hypothesis of efficient computation as an alternative for neural decoding, a memory-augmented Transformer whose feed-forward block is a low-rank instruction bank: a shared base operator plus a small set of learned low-rank instructions, from which a per-token code synthesizes the weight matrix ac...

Morteza Sarafyazd · 0 citations
#machine learning Review Sep 2026

Feature Superposition in Neural Networks: From Theory to Practice

This survey reviews the geometry, learning, and computation of superposed representations, explaining how feature statistics and decoder choice affect the conclusions and compares practical methods for recovering and analyzing features.

Dai Shi, Xiao-Yu Li, Andi Han et al. · 0 citations
Jul 2026

Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs

Prox is a two-stage training-free framework for sparse SwiGLU FFNs that outperforms training-free baselines at all sparsity levels, achieves up to a $1.99\times end-to-end decoding speedup at 70\% FFN sparsity, and is compatible with quantization and sparse attention.

Jinyi Liu, Wei Chen, Pengyu Chen et al. · 0 citations
Preprint Aug 2026

The Sparsity Whisperer

A family of difference-informed pruning methods built upon this principle are introduced, suggesting that preserving output differences is a broadly useful and composable signal for post-training LLM sparsification.

Linghao Kong, Inimai Subramanian, Micah Adler et al. · 0 citations
#machine learning Preprint Aug 2026

Sparse Competition during Training For the Emergence of Specialized Modules

This work introduces a method that maintains near-baseline accuracy, induces usage-based modularity by sparsely routing inputs to neuron groups, and encourages specialization of these modules, such that their activations are correlated with input classes.

Baptiste Rossigneux, Karim Haroun · 0 citations
#machine learning Preprint Sep 2026

Coarse-Graining Hidden Representations: Unsupervised Neuron Selection via Mapping Entropy

This work cast neuron selection as the problem of coarse-graining the hidden layer by retaining a subset of its neurons, and scores each putative selection by the mapping entropy (ME), measuring the loss of discriminatory power inherent in discarding part of the network neurons.

Margherita Mele, Andrea Castagna, R. Menichetti et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.