Collaborative Large Model Caching and Inference Offloading With Parameter Sharing in MEC
Pretrained Foundation Models (PFMs) enable highaccuracy inference services but are typically deployed in remote datacenters, resulting in prohibitively high inference delay. Mobile Edge Computing (MEC) can mitigate such high delays by caching PFMs or their fine-tuned variants on cloudlets located close to end users. Ho...