Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Representation Dynamics Reveal Semantic Saliency and Similarity for Visual Token Pruning in MLLMs

Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention maps or output features to estimate token importance or redundancy. Several recent approaches also exploit representation changes, but when and how these changes reflec...

Wei-Xuan Li, Zi-Kun Zhou, Xin-Yi Zhuang et al. · 0 citations
#machine learning Preprint Sep 2026

MoRA: MoE Pruning via Router Bias Learning and Expert Approximation

Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts....

Yu-Shuai Sun, Zi-Kun Zhou, Lin Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.