Skip to content
Conference

Carbon- and Cost-Aware Scheduling of Cloud AI Inference Workloads Across Geo-Distributed Regions

Aug 2026 · 2026 International Conference on Secure Information Systems and Technologies (ICSIST) · pp. 726-730 · 0 citations · 16 references

Abstract

Running artificial-intelligence inference in the cloud carries a rising energy and carbon cost, and where a request is served turns out to matter. Both the price of electricity and the carbon intensity of the local power grid differ from one region to another and change through the day. Most carbon-aware methods reduce emissions by delaying flexible batch jobs until the grid is cleaner, but an inference request cannot wait: it has to be answered within a tight time budget. This paper studies the spatial form of the problem—sending each latency-sensitive request to the region that lowers carbon and cost while still meeting its deadline. We express joint request routing and replica provisioning across geo-distributed regions as an online optimization that minimizes a weighted sum of operational carbon and monetary cost under latency and capacity limits, and we drive it with short-term forecasts of demand, price, and grid carbon intensity. A prototype built on Kubernetes with a v LLM inference backend, evaluated on daily regional profiles, lowers carbon by 32% and cost by 14% against nearest-region routing while keeping 99.2% of requests inside a 200 ms latency target.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.