Sep 2026· Proceedings of the International Conference on Parallel Processing· pp. 133-143· 0 citations· 16 references
Abstract
In modern GPU-based non-unified heterogeneous systems, CPU-GPU communication happens via the PCI bus. Data transfers are affected by startup overhead, which underutilizes the PCI channel bandwidth for small transfers. Modern programming models, such as CUDA and OpenMP, treat each input argument to a compute kernel independently, leading to data movement segmentation and execution slowdowns. This work presents In-Copy Fusion (ICF), a runtime optimization of OpenMP’s accelerator model, implemented in LLVM, that gathers eligible map clauses into staging buffers, orchestrating an optimal transfer pipeline through fusion without modifying application or kernel code. The technique preserves OpenMP semantics with negligible overhead. We evaluate ICF across a range of representative HPC benchmarks and configurations, including varying argument counts, data sizes, and argument-size disparities, as well as real-world benchmarks. The results show that ICF improves effective host-to-device (H2D) bandwidth and reduces end-to-end time relative to the baseline runtime with per-argument transfers across platforms and workloads, achieving an up to 4.8 × speedup.
This paper presents a zero-overhead, multi-vendor GPU compilation framework built natively into the Rust compiler (rustc) and LLVM backends, and leverages Rust's rich type system, ownership system, and strict aliasing guarantees to efficiently manage and optimize data transfers through LLVM's Offload infrastructure.
Manuel S. Drehwald, Marcelo Domínguez, Kevin Sala et al.· 1 citation
Modern LLM architectures and systems render GPU kernel input shapes increasingly dynamic, e.g., conditional expert routing in MoE and batches with variable-length sequences. This dynamism degrades performance in both expert-tuned kernels and compiler frameworks. We identify the root cause as dynamic cross-SM data depen...
Jing-Kai He, Guang-Da Sun, Tian-Jian Li et al.· Proceedings of the ACM SIGOP...· 0 citations
Moving from quantum research and development to production-grade, fault-tolerant quantum workload execution remains one of the most significant challenges facing quantum platform builders. While Python frameworks have enabled an easy entry point for quantum algorithm design, the low-latency requirements for real-time q...
Joseph K. L. Lee, M. Malekmohammadi, Hong-Sheng Zheng et al.· 0 citations
We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workl...
Genghan Zhang, Yixin Dong, Chengze Fan et al.· 1 citation
GPU-accelerated analytical query processing is often limited by both GPU device memory capacity and host-to-device data transfer time. Modern data compression techniques, such as cascaded lightweight compression, can mitigate these issues. However, existing designs all exhibit critical tradeoffs on compression ratios,...
Yong-Qi Zhuo, Xin-Yu Zeng, Huan-Chen Zhang et al.· Proceedings of the ACM on Ma...· 0 citations
As LLM inference scales and KV-cache pressure forces data onto NVMe storage, the frequency and fine-grained nature of GPU-initiated storage accesses have grown substantially. Existing GPU-initiated I/O platforms do not fully exploit NVMe’s multi-queue parallelism because their serialized submit-then-poll path keeps onl...
Jaewoong Jung, Hyukyul Kang, J. Choi· Proceedings of the 18th ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.