Skip to content
Open access

ORCHA: A performance portability system for extreme heterogeneity

Jul 2025 · The international journal of high performance computing applications · 1 citation · 34 references
Mathematics Computer Science

TL;DR

The core design goal is to let users decide the application’s hardware mapping and orchestration by editing only the high-level recipe—without modifying the maintained source code or binding the application to a particular runtime system.

Abstract

Heterogeneity is the prevalent trend in the rapidly evolving high-performance computing (HPC) landscape in both hardware and application software. The diversity in hardware platforms, currently comprising various accelerators and a future possibility of specializable chiplets, poses a significant challenge for scientific software developers aiming to harness optimal performance across different computing platforms while maintaining the quality of solutions when their applications are simultaneously growing more complex. Code synthesis and code generation can provide mechanisms to mitigate this challenge. We have developed a divide and conquer approach where different aspects of performance are handled by different stand-alone tools that are interfaced with the application through generated code. This portability system, ORCHA, enables users to configure and orchestrate their computations among available resources on a platform by specifying a high-level recipe, thereby permitting a many-to-many paradigm where each recipe results in a different variant of the application. The core design goal is to let users decide the application’s hardware mapping and orchestration by editing only the high-level recipe—without modifying the maintained source code or binding the application to a particular runtime system. Tools in ORCHA distribution are: CG-Kit for translating the recipe into an execution graph; Milhoja to execute the graph by orchestrating data and task movement among hardware resources; and Macroprocessor that enables users to define their own code-shorthand for higher composability and easier management of code variants. Additionally, the design of ORCHA permits tools to work in a plug-and-play mode where the application can build and run without CG-Kit and Milhoja, and either tool can be swapped out for other tools with similar capabilities by modifying the code generation portion of ORCHA. In this paper, we describe the design of ORCHA and the role that code-generation plays in isolating applications from tools. We demonstrate the breadth of configurations ORCHA enables with a case study in which an application configuration is realized on three distinct hardware mappings—a GPU-centric, a CPU/GPU balanced, and a CPU/GPU concurrent layouts by using different recipes.

Read PDF

Similar papers

On the co-design of runtimes, systems and programming interfaces for HPC

(English) High-performance computing (HPC) platforms are evolving towards increasingly complex architectures: many-core CPUs with multi-level NUMA hierarchies, heterogeneity with multiple classes of accelerators and higher-capacity interconnects. The increasing complexity and variety of resources in these machines make...

David Álvarez Robert · 0 citations
Preprint Sep 2026

Performance vs Portability in Heterogeneous HPC Environments: Why Pre-execution Benchmarking is Required

Cloud computing and high-performance computing (HPC) typically follow different paradigms: cloud services are often orchestrated using Kubernetes, whereas HPC workloads are managed through batch schedulers such as Slurm. Growing demand for shared computational resources increases the need for interoperability between t...

M. Mačernis · 0 citations
Book Open access Sep 2026

A SYCLic Investigation of SYCL Ecosystems: UniSYCL for Heterogeneous Scaling

Architectural diversity has turned accelerator performance portability into a compiler/runtime problem: portable source code is useful only if the surrounding ecosystem can also coordinate devices, backends, and data movement. This paper evaluates three SYCL ecosystems—Intel oneAPI DPC++, AdaptiveCpp, and UniSYCL—acros...

Nabayan Chaudhury, Norihisa Fujita, Beau Johnston et al. · 0 citations

Lessons Learned Building Cross-Architecture Analytical Engines

This thesis explores hardware-software co-design for data-intensive applications, targeting the unification of programming models using open standards and exploring experimental techniques for automated query synthesis, and presents X-BQSR, a holistic redesign of genomic base quality score recalibration pipelines.

I. D. Kabadzhov · 0 citations
Open access Oct 2026

When FPGA Meets Dataflow Analysis: An Explorative Step

Set-based (a.k.a. bit-vector-based) dataflow analysis is a fundamental building block for many static analysis tasks, and significant effort has been devoted to accelerating it. Existing acceleration approaches address the problem from a software perspective, leveraging various general-purpose computing platforms, such...

Fan-Juan Wei, Qin-Lin Chen, Nai-Ren Zhang et al. · 0 citations
Preprint Sep 2026

Compiler and Hardware Co-Design for Accelerator Architectures

Heterogeneous accelerator architectures offer an efficient path to performance for compute-intensive workloads. However, full-stack integration remains difficult. We present EAAC (Extensible Accelerator Architecture), a flexible and extensible compiler and hardware architecture designed to lower the overhead of hardwar...

Karl Herman Krause, Emad Jacob Maroun, Martin Schoeberl · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.