Aug 2026· Electronics· Vol 15, pp. 3522· 0 citations· 19 references
TL;DR
This paper proposes a compile-time instruction scheduling method that models sub-cycle resource usage and analyzes both data and structural dependencies at fine granularity, enabling precise detection of structural hazards in complex execution units.
Abstract
Application-specific instruction set processors (ASIPs) often employ specialized hardware to improve performance, but this introduces complexity in resource management and programming. Existing compiler solutions, including LLVM’s default schedulers, lack fine-grained structural conflict analysis for complex arithmetic logic unit (ALU) instructions, leading to suboptimal performance or runtime errors. This limitation becomes critical when targeting very long instruction word (VLIW) architectures with instruction fusion units that exhibit pipeline-stage-level resource contention. In this paper, we propose a compile-time instruction scheduling method that models sub-cycle resource usage and analyzes both data and structural dependencies at fine granularity. Unlike coarse-grained resource tables used in existing compilers, our approach tracks functional unit occupancy at the pipeline stage level, enabling precise detection of structural hazards in complex execution units. We implement this scheduler as a backend pass in the LLVM compiler framework and validate it on the Sayram VLIW processor for wireless communication. Experimental results show that our approach achieves 100% scheduling correctness while improving execution efficiency by 23% on average compared with in-order scheduling, with benefits up to 38% for highly parallel kernels such as PRACH, and reducing average running time by 66% compared with atomic execution.
Loom, a tuning-free symbolic compiler framework for tile-based SPMD programs on spatial dataflow architectures, is presented, suggesting that hardware-derived symbolic compilation provides a retargetable alternative to profiling-based tuning for spatial dataflow architectures while remaining interpretable by keeping op...
He-Ru Wang, Wei Li, Zhen-Yu Bai et al.· 0 citations
Training and serving large language models (LLMs) has become a core business for AI providers. To ensure a high-quality user experience while optimizing infrastructure costs, providers need to closely monitor the performance of LLM executions in production. However, existing performance profiling tools fall short in th...
Wei Liu, Yong-Chao He, Bo-Han Zhao et al.· Proceedings of the ACM SIGOP...· 0 citations
Dolunay is introduced, a RISC-V-based Independent Thread Scheduling (ITS) SIMT accelerator that employs a cooperative multitasking model and explicit synchronization barriers at the hardware-level, and provides the forward-progress guarantees necessary to implement starvation-free algorithms.
Ahmet Can, Erkan Uslu· WiPiEC Journal - Works in Pr...· 0 citations
This paper presents Ouros, a pipelined processor for lazy functional programming languages based on combinator graph reduction. Ouros overcomes the inherent sequentiality of graph reduction through dataflow-driven execution and automatic fine-grained multi-threading. It maintains pipeline utilisation by hiding per-thre...
Yu-Kang Xie, Craig R. Ramsay, Robert J. Stewart et al.· IEEE International Conferenc...· 0 citations
Fine-grained computation--communication overlap in distributed Mixture-of-Experts (MoE) inference allows communication to begin as partial compute results become ready. However, cooperative thread arrays (CTAs) performing computation and communication contend for finite residency capacity on streaming multiprocessors (...
LLM inference has become an essential service, yet it imposes unprecedented demands on memory bandwidth, computational density, and communication efficiency. While IMC is a promising solution to the memory wall issue, the heterogeneous data dynamicity of LLM requires complementary resources to handle intermediate data...