AutoUVM: Automated Prefetching Framework for LLMs under UVM Oversubscription
AutoUVM is proposed, an automated, framework-aware UVM prefetching system for efficient LLM execution under memory oversubscription and bridges the semantic gap between deep learning frameworks and UVM by exposing tensor-level access information and enabling policy-driven prefetching at fine granularity.