Enabling Spatially Fine-Grained DVFS in Neural Processing Units for Energy-Efficient LLM Serving
As neural processing units (NPUs) evolve rapidly to accommodate the ever-increasing compute demand of large language models (LLMs), their power consumption is becoming a limiting factor. Our study shows that using dynamic voltage and frequency scaling (DVFS) to exploit the service-level objective (SLO) slacks is a prom...