Skip to content

PowerScale: Energy-Efficient Geo-Distributed Model Training with Federated Datacenter Power

Jul 2026 · arXiv.org · Vol abs/2607.25650 · 0 citations · 42 references
Computer Science

TL;DR

PowerScale is presented, a hierarchical aggregation system that exploits the latency hierarchy of wide-area networks and shortens synchronization barriers and replaces per-site WAN transmissions with fewer, pre-aggregated transmissions at a lower frequency.

Abstract

The power demands of large-scale AI training increasingly exceed the capacity of any single data center, making geo-distributed training across power-constrained sites a practical necessity. Prior work optimizes such training mainly for time-to-accuracy using single-tier aggregation, where every site exchanges model updates directly with a central aggregator over the WAN each synchronization round, without accounting for the energy required to reach convergence. Single-tier aggregation is fundamentally energy-inefficient because synchronization barriers force faster sites to idle, full WAN updates dominate communication energy at scale, and fixed synchronization frequency keeps paying the same communication cost even when updates shrink late in training. To address these inefficiencies, we present PowerScale, a hierarchical aggregation system that exploits the latency hierarchy of wide-area networks. PowerScale organizes sites into regional clusters and applies a Sync-Async synchronization modality: sites synchronize frequently with a nearby cluster aggregator over fast local links, while cluster aggregators push pre-aggregated updates asynchronously to a global aggregator over the WAN. PowerScale forms clusters based on both network proximity and power availability, and uses an adaptive synchronization policy that reduces communication energy by adjusting how often clusters synchronize to training progress. This structure shortens synchronization barriers and replaces per-site WAN transmissions with fewer, pre-aggregated transmissions at a lower frequency. We evaluate PowerScale at 100-site scale in a Flower-based simulation environment. PowerScale matches or slightly improves time-to-accuracy compared with single-tier baselines while reducing energy consumption by up to 3.9x.

View source

Similar papers

Book Open access Aug 2026

GeoOrchestra: Orchestrating Heterogeneous Geo-Distributed Training with Network-Aware Scheduling

GeoOrchestra is a system that decouples resource filtering from fine-grained strategy search by abstracting compute nodes via computation and memory profiles while modeling WAN links as a virtual hard pipe, which employs hetero-aware pruning to filter invalid resource sets and a resource-driven search that exploits res...

Ting Liu, Qinghua Wu, Jun Zhou et al. · 0 citations
Preprint Sep 2026

Multi-Scale Datacenter Power Modulation

Cloud datacenters must increasingly modulate power in response to time-varying grid and infrastructure constraints. We study this problem as finite-horizon control of a networked hybrid dynamical system, where datacenter power and service capacity depend on interactions between servers, workers, and hosted services. Po...

Akshay Sreekumar, Nicolas H. Christianson, Fiodar Kazhamiaka et al. · 0 citations
Preprint Aug 2026

InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers

The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality. Yet operators often need to compare deployment alternatives before large-scale infrastructure is...

Nicoletta Tsiopani, Moysis Symeonides, G. Pallis et al. · 0 citations
Book Open access Aug 2026

Toward WAN-Aware LLM Training Across Heterogeneous, Geo-Distributed Sites

Preliminary results from a geo-distributed LLM training prototype that treats networking constraints as first-order design concerns are presented, motivating adaptive networking support for synchronization, compression, placement, telemetry, and recovery in geo-distributed LLM training.

Zi-Yue Luo, Jiaxuan Cai, Cedric Le Denmat et al. · 0 citations
Review Aug 2026

Slasher: Power Flexibility for Cloud Datacenters

Datacenters consume many megawatts of power, and regularly encounter scenarios that require modulating their power draw. These scenarios include datacenter infrastructure failures, power grid failures, grid services, and more, spanning a diverse range of requirements in terms of the power magnitude, the scope of the re...

Liuzixuan Lin, Fiodar Kazhamiaka, A. Kumbhare et al. · 1 citation
Book Open access Aug 2026

DistDPU: A Disaggregated DPU Architecture for High-Performance and Cost-Efficient AI Clouds

DistDPU is presented, a disaggregated DPU architecture that redefines the scaling abstraction for high-bandwidth cloud networking and co-designs the EM-OM functions and the inter-module fabric to minimize virtualization overhead while enforcing security and manageability invariants equivalent to those of a monolithic D...

Hao Mei, Lizhou Gao, Yuanyi Zhu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.