A carbon-aware routing framework that distributes function-calling queries across a three-tier edge-cloud architecture, combining edge and cloud LLMs on heterogeneous hardware and matches cloud-level accuracy while reducing operational carbon emissions by $4\times on average.
Abstract
Large Language Models (LLMs) with function-calling capabilities are becoming critical for modern agentic AI systems. Nevertheless, current deployments typically route inferences to powerful cloud-based models, incurring significant energy use and carbon emissions. We address this sustainability challenge with a carbon-aware routing framework that distributes function-calling queries across a three-tier edge-cloud architecture, combining edge and cloud LLMs on heterogeneous hardware. At its core, a lightweight k-NN predictor operating in a unified semantic-lexical embedding space estimates query-specific accuracy, delay, and power consumption on each edge tier. These predictions are then combined with real-time grid carbon intensity to route every query to the lowest-emission tier capable of executing it successfully. Evaluated on state-of-the-art function-calling benchmarks and LLM families, our framework matches cloud-level accuracy while reducing operational carbon emissions by $4\times$ on average.
An extensive critical review of serverless functions in cloud–edge environments reveals that cloud–edge serverless systems need accountable placement, state-aware workflows, reproducible benchmarking, trustworthy orchestration, and carbon-aware lifecycle control, which can be achieved only by going beyond latency and e...
Abdullah Abbasi, D. Hakro, Asad Ullah et al.· Future Internet· 0 citations
Deploying Large Language Models (LLMs) over the edge-cloud continuum faces severe stability challenges due to the conflict between stochastic network topology and complex workflow dependencies. Existing schedulers, relying either on computationally prohibitive Graph Neural Networks (GNNs) or topology-agnostic heuristic...
Yan Gao, Shaoyuan Huang, Yonghui Ye et al.· IEEE Transactions on Cogniti...· 0 citations
GreenBench, a benchmarking framework that evaluates the energy efficiency, throughput, and carbon footprint of five open-source LLMs across three NLP tasks on an Apple M4 Pro with 48 GB unified memory, is presented.
R. Kannan, Rajendra P. Firke, Shreya Bengle et al.· International Conference on...· 0 citations
Proliferation of cloud-based latency-sensitive workloads requires infrastructures tuned to their workload-specific latency constraints. Today, they shape the cloud from a generalized computing platform to diverse workload-specific cloud environments. As the demand for latency-sensitive workloads increases, cloud servic...
Tharindu B. Hewage, Shashikant Ilager, Maria Rodriguez Read et al.· 0 citations
This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers, using an LLM to predict key metrics such as execution time and energy consumption from source code.
Hanzhao Wang, Jingxuan Wu, Yumeng Li et al.· 0 citations
This work builds on Wang et al.'s taxonomy of Continuum Orchestration Systems employing DRL techniques and extends it with two further dimensions, measuring how LLMs are exploited and bridging the incommensurable per-tier signals and the LLM Orchestrator.
Antonino Vaccarella, Lan-Pei Li, Vincenzo Lomonaco et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.