The scale of data-intensive workloads in intelligent datacenters has grown rapidly in recent years, intensifying the need for efficient data movement within disaggregated memory pools across memory and storage subsystems. However, most existing solutions focus primarily on optimizing communication between compute nodes and data nodes, overlooking fine-grained management of in-pool data flows. Furthermore, prevailing data swap strategies often suffer from high overhead, limited parallelism, and a lack of adaptivity to dynamic workloads, especially for bursty data access patterns in heterogeneous memory pools. To address these issues, this paper introduces MEDO, a high-parallelism data offloading system for disaggregated memory pools. MEDO leverages a novel multi-stream data offloading architecture, featuring parallel data streams and approximate LRU queues, to maximize throughput and efficiently handle diverse workloads. Additionally, MEDO incorporates a lightweight, adaptive offloading agent that dynamically optimizes data placement decisions and fine-grained system configurations. Our prototype achieves up to 3.6 × latency reduction on real-world data services compared with baselines and can reduce in-pool memory usage by up to 50% on state-of-the-art disaggregated memory systems.
Jing Wang, Han-Zhang Yang, Chao Li et al.· Proceedings of the Internati...· 0 citations
The increasing use of renewable energy in data centers creates an opportunity to reduce the carbon footprint of energy-intensive LLM inference workloads. Unlike traditional stable power supply, renewable generation fluctuates over time, making it difficult to match computation demand with available energy. However, existing LLM serving systems primarily optimize latency and throughput without considering energy supply dynamics, leading to underutilization of renewable energy and unnecessary reliance on thermal power, and consequently, higher carbon emissions. We present GreenAlign, a renewable-aware scheduling framework that addresses this mismatch by treating best-effort (BE) requests as temporally shiftable load. GreenAlign enforces a power-constrained policy that executes BE requests using only residual renewable energy under normal conditions, and introduces a backlog risk metric to selectively relax this constraint when deadline violations are imminent. To ensure responsiveness, it maintains standby capacity to absorb unpredictable latencycritical (LC) bursts and uses lightweight length estimation to handle request uncertainty. Simulation results show that GreenAlign significantly reduces thermal energy usage while preserving LC latency and BE deadline satisfaction.
Chang Liu, Jiacheng Liu, Xiaofeng Hou et al.· Fall Joint Computer Conferen...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.