Skip to content

Rethinking LLM-aided RTL Code Optimization Via Timing Logic Metamorphosis

Sep 2026 · ACM Transactions on Reconfigurable Technology and Systems · 0 citations · 66 references

TL;DR

A metamorphic testing based evaluation method and a companion benchmark to measure the effectiveness of LLM-aided RTL code optimization show that LLM-aided methods effectively optimize logic operations and data paths, achieving lower or stable wire counts and area with delays near baseline.

Abstract

Register Transfer Level (RTL) code optimization is critical for meeting performance and power budgets in Field Programmable Gate Array (FPGA) design. Traditional approaches depend on expert knowledge or heuristics, which are time consuming and error prone. Therefore, recent research has explored using Large Language Models (LLMs) for RTL code optimization and providing initial evidence of their potential. However, the capabilities and limitations of LLMs in RTL code optimization have not been systematically assessed, especially for RTL with complex timing logic. Moreover, these studies cannot distinguish whether improvements come from genuine reasoning or merely reproducing memorized patterns from training data. To address these problems, we present a metamorphic testing based evaluation method and a companion benchmark to measure the effectiveness of LLM-aided RTL code optimization. Our key idea is that optimization effectiveness should be consistent across RTL descriptions that are semantically equivalent. We first build a benchmark suite covering four domains (logic operations, datapaths, timing control flow, and clock domains). We then apply domain specific transformations to produce semantically equivalent RTL code with different code forms. We assess each method by applying it to both the original and the transformed code, synthesizing them with the same flow, and comparing the synthesis results. Through extensive experiments, our results show that LLM-aided methods effectively optimize logic operations and data paths, achieving lower or stable wire counts and area with delays near baseline. For timing control flow and clock domains the gains are limited and structural ratios often rise. Larger code size also increases the chance of incorrect rewrites, especially in timing and cross domain designs. We further explore the application of prompt engineering include few shot examples and Chain of Thought (CoT) to improve RTL code optimization. Based on these results we discuss practical directions and give concrete suggestions for using LLMs in RTL optimization.

View source

Similar papers

Preprint Sep 2026

GRADE-RTL: Evaluating LLM-Generated RTL Beyond Compilation

Large language models (LLMs) can generate register-transfer-level (RTL) code from natural-language specifications, but compilation alone does not establish structural completeness, functional correctness, or implementation efficiency. This paper presents a framework for evaluating LLM-generated RTL beyond compilation,...

Hepziba Susan, R. ShivaranjaniG, Malik Imran et al. · 0 citations
Open access Aug 2026

SynaSpace: Behavior-Driven Configuration Optimization of Test Generators for Logic Synthesis Testing

SynaSpace is a behavior-driven configuration optimization framework for fault detection in logic synthesis tools that focuses on synthesis behavior coverage to guide configuration search, by constructing behavioral representations through joint analysis of synthesis logs and gate-level netlists.

Pei-Yu Zou, Xiao-Chen Li, Yijia Meng et al. · 0 citations
Preprint Aug 2026

SoK: ARCUS: On the Efficiency and Efficacy of Hardware Fuzzing

This work presents a comprehensive analysis of contemporary hardware fuzzing techniques applied across three major abstraction layers: Instruction Set Architecture (ISA), microarchitecture, and Register-Transfer Level (RTL). Our study examines key factors including input stimulus quality, mutation strategies, feedback...

Alenkruth Krishnan Murali, Raghul Saravanan, D. SaiManojP et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PerfReasoning: How Well Do LLMs Reason on Hardware Performance?

Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical...

Da Zhao, K. Sankaralingam, Christos Kozyrakis et al. · 0 citations
Preprint Aug 2026

Effect of Abstractions and Prompting Strategies on LLM-Guided High-Performance Optimizations

It is demonstrated that LLMs provided with specific optimization goals achieve better measured performance and validity rates when generating C code compared to creating computation pipelines and optimization schedules with established frameworks, suggesting that future development should explore alternative approaches...

Jiří Klepl, Matyáš Brabec, Martin Kruliš · 0 citations
Preprint Aug 2026

T-LLM Compiler: Trusted LLM-based Code Optimization and Verification Framework

The Trusted LLM (T-LLM) Compiler is presented, which proposes an advancement in compiler technology through a collaborative effort involving high-level LLM code transformations, traditional compilers, and verification tools and facilitates iterative code optimization efforts with verification strategies that enable cor...

Zahra Fazel, Sunanda Gamage, Shayan Shirahmad Gale Bagi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.