Edge platforms increasingly rely on GPUs and deep learning accelerators (DLAs) to support real-time computing of emerging workloads. However, their power, frequency, and security characteristics remain largely unexplored. This paper provides the first in-depth characterization of power and frequency behaviors of NVIDIA...
M. Rafi, Kevin Chau, Hyeran Jeon· ACM Transactions on Architec...· 0 citations
Large language model (LLM) outputs are expected to be reproducible under greedy decoding, yet in practice the same model, prompt, and software stack produce different outputs on different GPUs. The root cause is floating-point non-associativity combined with hardware-dependent kernel selection. Inference frameworks sel...
L. Cooper, Shinnung Jeong, Hyeran Jeon et al.· 0 citations
AutoUVM is proposed, an automated, framework-aware UVM prefetching system for efficient LLM execution under memory oversubscription and bridges the semantic gap between deep learning frameworks and UVM by exposing tensor-level access information and enabling policy-driven prefetching at fine granularity.
Mao Lin, Hui Feng, Xian-Zhong Ding et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.