CC-Bench is presented, a lightweight, extensible, and application-oriented benchmark suite for evaluating communication compression under realistic execution conditions, and uses declarative application-environment modeling to decouple profiling logic from communication libraries, datasets, and fidelity metrics, enabling portable cross-library evaluation.
Abstract
Distributed HPC and LLM workloads increasingly require efficient communication for scalability, yet growing data movement has become a major performance bottleneck. Communication compression can reduce this overhead and complement execution-level optimizations, but its benefits remain difficult to assess because existing benchmarks lack support for diverse backends, realistic datasets, application-specific accuracy metrics, and overlap-induced resource contention. We present CC-Bench, a lightweight, extensible, and application-oriented benchmark suite for evaluating communication compression under realistic execution conditions. CC-Bench uses declarative application-environment modeling to decouple profiling logic from communication libraries, datasets, and fidelity metrics, enabling portable cross-library evaluation. It further combines function-level interception and hardware counter monitoring to characterize per-phase latency, hardware utilization, numerical fidelity, and computation interference. With representative datasets from HPC and LLM workloads, CC-Bench evaluates three compression-enabled communication libraries on CPU and GPU clusters, revealing accuracy-performance trade-offs and bottlenecks to guide practical deployment and optimization.
Cloud computing and high-performance computing (HPC) typically follow different paradigms: cloud services are often orchestrated using Kubernetes, whereas HPC workloads are managed through batch schedulers such as Slurm. Growing demand for shared computational resources increases the need for interoperability between t...
This paper compares three parallel programming paradigms—shared memory (OpenMP), distributed memory (MPI), and heterogeneous GPU computing (CUDA)—for a telemetry-inspired Map-Filter-Reduce-Sort analytics pipeline. The study evaluates datasets ranging from 100 million to 1 billion elements and focuses on end-to-end beha...
Krystian Ochmański, Filip Krużel, M. Nytko· IEEE Access· 0 citations
Characterising AI workload performance on modern HPC systems requires understanding both their scalability in isolation and their behaviour under concurrent execution. However, the interplay among parallelisation strategies, network congestion, compute capability, and interconnect technologies remains poorly understood...
Jacopo Raffi, Thomas Pasquali, Lorenzo Piarulli et al.· 0 citations
We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workl...
Genghan Zhang, Yixin Dong, Chengze Fan et al.· 1 citation
(English) High-performance computing (HPC) platforms are evolving towards increasingly complex architectures: many-core CPUs with multi-level NUMA hierarchies, heterogeneity with multiple classes of accelerators and higher-capacity interconnects. The increasing complexity and variety of resources in these machines make...
David Álvarez Robert· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.