MoEP: Compact and efficient sparsity with modular expert paths.
The transition from dense to sparse model architectures has become a key trend in the field of Large Language Models (LLMs). Mixture-of-Experts (MoE) methods can be used to increase model conditional representation capacity by activating only a subset of parameters for each input token. However, the practical efficienc...