GPU and CPU Memory Co-Optimization in Heterogeneous Pipeline Parallelism for Efficient Large Language Model Fine-Tuning on Commodity Servers
Tiny-Pipe comprises a holistic layer packing method that simultaneously reduces GPU memory footprint and improves training performance, an active CPU memory management that alleviates CPU memory pressure by eliminating redundant parameters, and a layer-wise runtime swapping strategy that further enhances overall performance.