Skip to content

CausalGate: Causal Importance Distillation for Transformer Module Pruning

Jul 2026 · arXiv.org · Vol abs/2607.22720 · 0 citations · 29 references
Computer Science Mathematics

TL;DR

This work introduces CausalGate, an intervention-guided framework for compute-efficient transformer inference, which consistently outperforms prominent dynamic routing and layer-skipping baselines, translating theoretical compute savings into concrete hardware latency reductions with zero operational overhead.

Abstract

Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitudes, to drop redundant modules. However, these correlation-based metrics often fail to capture subtle, non-linear structural computations vital for semantic accuracy. We introduce CausalGate, an intervention-guided framework for compute-efficient transformer inference. During a calibration phase, CausalGate isolates individual Attention and MLP sub-layers, zeros out their respective outputs, and measures the exact semantic damage via the Kullback-Leibler divergence of the final logit distribution. To eliminate runtime routing overhead, this structural importance hierarchy is distilled into a global set of static, lightweight scalar gates using an Exponential Moving Average smoothing objective paired with a differentiable pairwise ranking loss. Evaluated on TinyLlama-1.1B, Qwen2.5-3B, and Llama-3.1-8B across language modeling and commonsense reasoning benchmarks, CausalGate consistently outperforms prominent dynamic routing and layer-skipping baselines, translating theoretical compute savings into concrete hardware latency reductions with zero operational overhead.

View source

Similar papers

Preprint Aug 2026

RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models

It is found that preserving clean routing substantially recovers steering performance and that routing is more sensitive to semantic content than to behavioral changes under controlled content, which supports routing consistency as an important architectural consideration for adapting representation engineering to MoE...

Zhi-Bo Zhang, Zheng-Mao Ouyang, Ling Shi et al. · 0 citations
Preprint Aug 2026

DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference

This work proposes DistillCache, a reinforcement learning framework that formulates KV-cache eviction as a sequential decision problem and demonstrates the effectiveness of learned, distribution-aware policies for memory-efficient long-context LLM inference.

Asaad Althoubi · 0 citations
Jul 2026

Grounding latent algorithm routing in transformer reasoning

A central question in the in-context learning literature is whether transformers can organize episode-level adaptation around different inductive-bias families. We study this question in a controlled setting through latent algorithm routing: route-like behavior in which the solver-family preference changes with the lat...

Xiangbo Zhang, Xiaoxu Ma · 1 citation
Preprint Aug 2026

Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality

ADU is presented, a fine-grained, training-based framework that shifts unlearning from token erasure to contextual attention-pathway decoupling, and achieves the strongest aggregate performance among evaluated baselines on the TOFU and WMDP benchmarks.

Xun-Lei Chen, Qirui Ye, Yuang Li et al. · 0 citations
Preprint Aug 2026

LegoLM: Structured Weight Sharing for Large Language Models

It is discovered that outlier dominance grows with model scale: full replacement at K=128 degrades GPT-2 small but catastrophically degrades Mistral-7B by +1,134,279%, while selective replacement at p=99% rescues both models to under +15%.

Joseph Bingham · 0 citations
Preprint Aug 2026

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

GradCuit (gradient through circuit), which inserts optimizable latent states at a selected Transformer layer between the hidden representations of the prompt and the generated continuation, opens a new axis of robust and interpretable test-time scaling, where LLMs adapt how they reason rather than merely regenerate, sa...

Zhaoxin Yu, Qianli Shen, Hengli Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.