Skip to content
Book

Analyzing HPC Job Wait Times under Resource Scaling Using Historical Workload Data

Jul 2026 · Practice and Experience in Advanced Research Computing · 0 citations · 12 references
Computer Science

TL;DR

This work presents a data-driven framework that leverages historical job traces to estimate the impact of resource modifications on queue performance, and introduces the Weighted Wait-Time Score (WWS), a bounded metric that captures both typical and tail wait-time behavior.

Abstract

Understanding the impact of hardware configuration, infrastructure investments, and operational policies on job wait times in high-performance computing (HPC) systems is a challenging problem primarily due to the lack of effective tools that are built on controlled, real-world observations across different system configurations. In this work, we present a data-driven framework that leverages historical job traces to estimate the impact of resource modifications on queue performance. Using job data collected from multiple HPC systems during periods with no hardware changes, we construct a synthetic scheduler that replays real job submission patterns and resource requests, while using ground-truth runtimes to model execution. This enables controlled, counterfactual evaluation of infrastructure changes without modifying production systems. We perform experiments by scaling system resources, including GPUs and CPU cores. Our results on GPU-based systems show that increasing GPU capacity leads to significant reductions in wait times, while CPU-only scaling provides minimal benefit. To quantify these effects, we introduce the Weighted Wait-Time Score (WWS), a bounded metric that captures both typical and tail wait-time behavior. Our experiments show that informed resource scaling improves WWS (up to 92% relative gain in WWS for low-baseline systems), capturing gains in both typical and tail wait-time behavior under this metric. We further formulate a cost-aware optimization framework to guide resource allocation under budget constraints. Our approach offers a data-driven way to evaluate HPC upgrades and supports future predictive optimization and better resource use. Overall, this work provides (i) a trace-driven framework for evaluating resource scaling effects, (ii) a bounded metric (WWS) for comparing wait-time performance, and (iii) a pathway to optimal user wait time optimization based resource allocation in HPC systems.

View source

Similar papers

Book Open access Aug 2026

Theseus: Runtime-Adaptive GPU Collective Communication with Hot-Swappable Schedules

Current GPU Collective Communication Libraries (CCLs) employ predefined schedules optimized for stable environments. Their supported schedules and selection logic are fixed at communicator initialization, which fails to account for evolving runtime conditions, such as workload characteristics and hardware health status. Consequently, long-running GPU jobs experience suboptimal performance after hours or days of execution, which translates into longer job completion times and wasted GPU cluster resources. To address this problem, we present Theseus, a novel CCL backend that provides schedule-level runtime adaptivity. It admits user-defined schedules and selection policies. As runtime conditions change, Theseus selects suitable schedules using cluster-wide runtime attributes beyond CCL-internal metrics. Moreover, it hot-swaps from the previous schedule consistently across GPUs with low overhead. Theseus acts as a drop-in replacement to facilitate integration. We evaluate Theseus extensively on various GPU workloads with intuitive policies. Compared with NCCL, Theseus achieves up to 1.61X speedup of communication time in stable environments and 2.46X in dynamic environments. It improves end-to-end job completion time by up to 1.84X while incurring comparable or lower overhead.

Rui Ding, Xiandong Lu, Jiajun Wang et al. · 0 citations
Open access Aug 2026

Benchmarking Python–Rust Integration for High-Performance Computing Tasks

High-performance computing (HPC) requirementshave dramatically increased in recent years for a wide range oftasks, including data analytics, machine learning and system optimization. Although Python is a favored language in science and analysis, its dynamic features can cause performance bottlenecks.On the other hand, Rust is a new systems language that guarantees memory safety while offering near-native execution performance, and as such is a prime contender for high-performanceapplications. In this paper we conduct a benchmarking study of Python-Rust interoperability: We benchmark pure Pythonversions against Python code with embedded Rust routines (usingthe PyO3 package) for various tasks like arithmetic calculations,string operations, list operations, file operations and conditionalstatements. Our benchmarking experiments reveal that the Rust enhanced versions systematically execute faster than pure Pythoncode- up to three times faster in some cases. We also analysethe impact of effective language integration and reduced run-timeon more sustainable software engineering practices: by reducingthe overall number of CPU cycles, memory use and energy consumption, hybrid approaches enable more energy-efficient HPCapplications. Finally, we discuss when and how the Python–Rust co-development can be used to create high-performance andenergy-efficient solutions in scientific computing.

Srikant Singh, Ravi Raj · 0 citations

On the co-design of runtimes, systems and programming interfaces for HPC

(English) High-performance computing (HPC) platforms are evolving towards increasingly complex architectures: many-core CPUs with multi-level NUMA hierarchies, heterogeneity with multiple classes of accelerators and higher-capacity interconnects. The increasing complexity and variety of resources in these machines make it harder for application programmers to use them efficiently and effectively. As a result, many resources in modern HPC clusters remain underutilized, limiting energy efficiency and the potential throughput of the machine. This thesis argues that addressing these challenges requires co-design across the software stack, from runtime mechanisms and programming interfaces to system-level policies. We first study task-based runtimes and identify opportunities to reduce overheads and improve scalability on many-core machines, introducing novel scheduling and dependency-management techniques that maintain throughput under extreme concurrency. We then tackle programmability and performance in heterogeneous systems, proposing runtime and interface support to better overlap data movement, accelerator offloading, and CPU computation. We also make a case for the effective co-design of applications and programming models through the study of an increasingly common application class: iterative data-flow computations, commonly used in simulations, iterative solvers, and AI. Through this study, we propose specific optimizations for this application class, showing how holistic co-design approaches can lead to significant speedups. Finally, we present the nOS-V library with the goal of improving system-wide utilization through application co-scheduling, and we later apply the same methodology to obtain truly interoperable programming models, enabling multiple runtimes and parallel libraries to coexist within a single application with reduced mutual interference. Overall, the contributions of this thesis provide a set of runtime techniques, programming abstractions, and system mechanisms that jointly improve throughput, efficiency, and composability on next-generation HPC systems. (Català) Les plataformes de Computació d'Altes Prestacions (HPC) estan evolucionant cap a arquitectures cada cop més complexes: processadors amb grans quantitats de nuclis i amb jerarquies de memòria d'accés no uniforme (NUMA) multi-nivell, heterogeneïtat en els sistemes amb múltiples tipus d'acceleradors, i interconnexions de més capacitat. L'increment de complexitat i la varietat de recursos en aquestes màquines complica la feina d'utilitzar-les de forma eficient i efectiva, la qual recau en les persones que les programen. El resultat és que, als centres de dades moderns, molts recursos acaben sent infrautilitzats, empitjorant la seva eficiència energètica i limitant la seva capacitat. Aquesta tesi proposa abordar aquests reptes mitjançant el disseny conjunt de components a totes les capes del programari, anant des dels mecanismes d'execució i les interfícies de programació fins al programari de sistema de baix nivell. Primerament, estudiem els mecanismes d'execució basats en tasques i identifiquem oportunitats per reduir el cost computacional associat a la gestió de tasques i millorar l'escalabilitat en sistemes amb molts nuclis, introduint noves tècniques de planificació i gestió de dependències que mantenen el rendiment fins i tot en nivells extrems de concurrència. A continuació, abordem els problemes de programabilitat i rendiment en sistemes heterogenis, proposant noves interfícies i sistemes d'execució per superposar de manera més efectiva el moviment de dades, l'execució en acceleradors i el càlcul al processador principal. També argumentem a favor d'un disseny conjunt entre aplicacions i models de programació mitjançant l'estudi d'una classe d'aplicacions cada vegada més habituals: els càlculs iteratius basats en dependències de dades, utilitzats sovint en simulacions, mètodes matemàtics iteratius i intel·ligència artificial. A partir d'aquest estudi, proposem optimitzacions específiques per a aquesta classe d'aplicacions, mostrant com el disseny conjunt pot conduir a millores de rendiment significatives. Finalment, presentem el programari de gestió de tasques nOS-V amb l'objectiu de millorar la utilització global del sistema mitjançant la coplanificació d'aplicacions, i posteriorment apliquem la mateixa metodologia per aconseguir models de programació realment interoperables, permetent que diversos models i biblioteques paral·leles convisquin dins d'una mateixa aplicació minimitzant la interferència mútua. En conjunt, les contribucions d'aquesta tesi proporcionen tècniques d'execució, abstraccions de programació i mecanismes de sistema que, de manera conjunta, milloren el rendiment global, l'eficiència i la composabilitat dels sistemes HPC de nova generació. (Español) Las plataformas de Computación de Altas Prestaciones (HPC) están evolucionando hacia arquitecturas cada vez más complejas: procesadores con grandes cantidades de núcleos y con jerarquías de memoria de acceso no uniforme (NUMA) multinivel, heterogeneidad en los sistemas con múltiples tipos de aceleradores, e interconexiones de mayor capacidad. El incremento de complejidad y la variedad de recursos en estas máquinas dificultan su uso eficiente y efectivo, una tarea que recae en las personas que las programan. Como resultado, en los centros de datos modernos muchos recursos terminan infrautilizados, empeorando su eficiencia energética y limitando su capacidad. Esta tesis propone abordar estos retos mediante el diseño conjunto de componentes en todas las capas del software, abarcando desde los mecanismos de ejecución y las interfaces de programación hasta el software de sistema de bajo nivel. En primer lugar, estudiamos los mecanismos de ejecución basados en tareas e identificamos oportunidades para reducir el coste computacional asociado a la gestión de tareas y mejorar la escalabilidad en sistemas con muchos núcleos, introduciendo nuevas técnicas de planificación y gestión de dependencias que mantienen el rendimiento incluso en niveles extremos de concurrencia. A continuación, abordamos los problemas de programabilidad y rendimiento en sistemas heterogéneos, proponiendo nuevas interfaces y sistemas de ejecución para superponer de manera más efectiva el movimiento de datos, la ejecución en aceleradores y el cálculo en el procesador principal. También defendemos un diseño conjunto entre aplicaciones y modelos de programación mediante el estudio de una clase de aplicaciones cada vez más habituales: los cálculos iterativos basados en dependencias de datos, utilizados frecuentemente en simulaciones, métodos matemáticos iterativos e inteligencia artificial. A partir de este estudio, proponemos optimizaciones específicas para esta clase de aplicaciones, mostrando cómo el diseño conjunto puede conducir a mejoras significativas de rendimiento. Finalmente, presentamos el software de gestión de tareas nOS-V con el objetivo de mejorar la utilización global del sistema mediante la coplanificación de aplicaciones, y posteriormente aplicamos la misma metodología para lograr modelos de programación verdaderamente interoperables, permitiendo que diversos modelos y bibliotecas paralelas coexistan dentro de una misma aplicación minimizando la interferencia mutua. En conjunto, las contribuciones de esta tesis proporcionan técnicas de ejecución, abstracciones de programación y mecanismos de sistema que, de manera conjunta, mejoran el rendimiento global, la eficiencia y la componibilidad de los sistemas HPC de nueva generación.

Unknown authors · 0 citations
Book Open access Jul 2026

Extending the Life of HPC Systems in Resource Constrained Environments: Mapping Productivity-Energy Trade-offs in Memory-Bound Workloads via DVFS and Core Scaling on Repurposed Hardware

Many resource-constrained environments rely on repurposed hardware for High-Performance Computing, shifting costs from capital expenditure to operational energy costs. As a result, evaluating system viability in resource-constrained environments warrants a shift from measuring raw performance to evaluating Productivity-to-Energy efficiency. Memory-bound workloads are particularly impacted by the memory wall, where stalled processors waste energy. While Dynamic Voltage and Frequency Scaling is widely used, systematic core scaling remains largely overlooked. This work examines the impact of limiting active cores on repurposed nodes. Presenting Phase 1 preliminary results, initial HPCG benchmarking demonstrates that targeted core deactivation yields a 62.5% improvement in PTE efficiency over maximum-performance baselines. To address viability in complex applications such as OpenFOAM, a second phase introduces deep C-state power-gating, fully saturated workloads, and hardware-level power measurements. The resulting framework provides a practical, software-driven approach to lowering Total Cost of Ownership and advancing sustainable High-Performance Computing in resource-constrained settings.

Bryan Johnston, Suné Toerien, Vele Nefale et al. · 0 citations
Preprint Aug 2026

BOOSTEDSOSA: Accelerated Inferencing for Low Variance Stochastic Online Scheduling

Heterogeneous scheduling in stochastic, online envi- ronments, such as high-performance computing (HPC) systems, presents a significant challenge. Stochastic Online Scheduling Accelerators (SOSAs) offer a promising solution, but their effectiveness is compromised by a reliance on runtime estimates provided by users. These estimates introduce substantial vari- ance into the scheduling process (mean MAE in hundreds of Core-Days), thereby weakening the competitiveness of Stochastic Online Scheduling algorithms as their competitive-ratio bound increases with runtime variability. To address this limitation, we introduce BOOSTEDSOSA, a dual-FPGA ML-assisted Scheduling architecture that integrates a Machine Learning predictor for expected processing times, with a novel temporal-aware training policy. The predictor estimates job runtimes using only scheduler parameters available at submission time, enabling its use in existing HPC systems. Using historical real-world HPC job data (from the Argonne Leadership Comput- ing Facility, MIT Supercloud and UIUC Blue Waters workload datasets), we show that the predictor reduces MAE by up to 63.85% compared to user runtime estimates, and the additive training policy reduces MAE by up to 71.88% compared to a static model. End-to-end, BOOSTEDSOSA achieves an average 17x speedup over an AVX-optimized software baseline and processes up to 1,711 jobs/seconds

Adam H. Ross, Riccardo Revalor, Aryan Singh et al. · 0 citations