Skip to content

Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems

Sep 2026 · 1 citation · 19 references
Computer Science

TL;DR

A carbon-aware routing framework that distributes function-calling queries across a three-tier edge-cloud architecture, combining edge and cloud LLMs on heterogeneous hardware and matches cloud-level accuracy while reducing operational carbon emissions by $4\times on average.

Abstract

Large Language Models (LLMs) with function-calling capabilities are becoming critical for modern agentic AI systems. Nevertheless, current deployments typically route inferences to powerful cloud-based models, incurring significant energy use and carbon emissions. We address this sustainability challenge with a carbon-aware routing framework that distributes function-calling queries across a three-tier edge-cloud architecture, combining edge and cloud LLMs on heterogeneous hardware. At its core, a lightweight k-NN predictor operating in a unified semantic-lexical embedding space estimates query-specific accuracy, delay, and power consumption on each edge tier. These predictions are then combined with real-time grid carbon intensity to route every query to the lowest-emission tier capable of executing it successfully. Evaluated on state-of-the-art function-calling benchmarks and LLM families, our framework matches cloud-level accuracy while reducing operational carbon emissions by $4\times$ on average.

View source

Similar papers

Review Open access Sep 2026

Serverless Functions in Cloud–Edge Environments: A Comprehensive Critical Review and Taxonomy

An extensive critical review of serverless functions in cloud–edge environments reveals that cloud–edge serverless systems need accountable placement, state-aware workflows, reproducible benchmarking, trustworthy orchestration, and carbon-aware lifecycle control, which can be achieved only by going beyond latency and e...

Abdullah Abbasi, D. Hakro, Asad Ullah et al. · 0 citations
2026

Workflow-Aware Expert Routing for Distributed LLM Serving Over the Edge-Cloud Continuum

Deploying Large Language Models (LLMs) over the edge-cloud continuum faces severe stability challenges due to the conflict between stochastic network topology and complex workflow dependencies. Existing schedulers, relying either on computationally prohibitive Graph Neural Networks (GNNs) or topology-agnostic heuristic...

Yan Gao, Shaoyuan Huang, Yonghui Ye et al. · 0 citations
#artificial intelligence Conference Open access Aug 2026

Greenbench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source Llm Inference on Apple Silicon

GreenBench, a benchmarking framework that evaluates the energy efficiency, throughput, and carbon footprint of five open-source LLMs across three NLP tasks on an Apple M4 Pro with 48 GB unified memory, is presented.

R. Kannan, Rajendra P. Firke, Shreya Bengle et al. · 0 citations
Preprint Sep 2026

Carbon-aware Resource Management for Latency-Sensitive Cloud Computing Environments: A Taxonomy and Future Directions

Proliferation of cloud-based latency-sensitive workloads requires infrastructures tuned to their workload-specific latency constraints. Today, they shape the cloud from a generalized computing platform to diverse workload-specific cloud environments. As the demand for latency-sensitive workloads increases, cloud servic...

Tharindu B. Hewage, Shashikant Ilager, Maria Rodriguez Read et al. · 0 citations
Preprint Aug 2026

LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations

This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers, using an LLM to predict key metrics such as execution time and energy consumption from source code.

Hanzhao Wang, Jingxuan Wu, Yumeng Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Smart Adaptive Computing Across the Continuum: LLMs in IoT-Edge-Cloud Resource Management

This work builds on Wang et al.'s taxonomy of Continuum Orchestration Systems employing DRL techniques and extends it with two further dimensions, measuring how LLMs are exploited and bridging the incommensurable per-tier signals and the LLM Orchestrator.

Antonino Vaccarella, Lan-Pei Li, Vincenzo Lomonaco et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.