DejaVu: Unifying Memory Allocations to Eliminate Redundant Copies on Unified-Memory SoCs
GPU applications on unified-memory (UMA) edge platforms often inherit a discrete-GPU memory abstraction in which they allocate one buffer for the CPU, another for the GPU, and copy data between them before and after GPU execution. On UMA hardware these buffers reside in the same physical DRAM, so the copies consume ban...