Skip to content

SpecRead: A Benchmark for Measuring Whether Language Models Understand Hardware Specifications

Sep 2026 · 0 citations · 23 references
Computer Science

TL;DR

This work presents SpecRead, a benchmark that isolates specification comprehension from generation ability, and is automatically scorable by deterministic checks, with gray-zone cases counted wrong under the conservative main scoring.

Abstract

Existing benchmarks for large language models (LLMs) in hardware design evaluate downstream artifacts such as generated RTL, assertions, or testbenches. When a model fails such a benchmark, the failure is ambiguous: it may have misread the specification, or it may have understood the specification and failed to write the code. We present SpecRead, a benchmark that isolates specification comprehension from generation ability. SpecRead v2.1 contains 385 questions over 10 open-source OpenTitan IP blocks: exact retrieval, cross-section reasoning, contradiction detection in mutated specifications, and spec-RTL consistency checking, plus 82 controls (41 distractor, 41 consistent-RTL). Type-4 items are built from real RTL mutations; we retain only mutations that Icarus Verilog simulation shows to change observable behavior. A with-spec vs. without-spec ablation suggests the questions require the excerpt, not training recall alone (without-spec accuracy 3/20 on the t1/t2 subset), though memorization of the source text may still help spot mutations. As an initial characterization with a small model, Ministral-3B scores 33.2% overall (128/385; macro average 39.0%): 55.2% on retrieval, 51.7% on cross-section reasoning. On the two contradiction-focused types, the verdict-plus-location measure gives 48.0% (t3) and 63.3% (t4), with a 51.2% false-positive rate on distractors and 100% on consistent-RTL controls. Layered scoring shows the model locates contradictions well (78.9-81.6% location accuracy) but scores lower on their category (43.9-49.7%). A structured"rule-table"prompting intervention lowers accuracy on every question type except t2 (tied). SpecRead is automatically scorable by deterministic checks, with gray-zone cases counted wrong under the conservative main scoring. The benchmark is regenerable for type-3 items via mutation injection, and built exclusively from public sources.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Self-Spec Verifiable Code Generation

Large language models (LLMs) may generate unreliable code on corner cases missed by testing, while formal verification can provide machine-checkable guarantees. Recently, researchers have proposed several benchmarks to evaluate the capabilities of LLMs in generating formally verifiable code, where LLMs need to formulat...

Jia-Ru Qian, Yihong Dong, Yong-Ming Li et al. · 0 citations
Open access Aug 2026

How well do LLMs understand code?

SemBench is introduced, a novel benchmark consisting of 1000 diverse C programs sourced from the CodeParrot GitHub-code dataset, with 15,404 semantic questions spanning six basic but fundamental properties: dead code-statement, data dependency, function reachability, dominator, dead code-loop, and liveness.

Jade Xu, Ren-Liang Sun, Zijian Ding et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

SWE-Flux is introduced, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written manually or judged by LLMs.

Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband et al. · 0 citations
#small language model Preprint Sep 2026

Path2Spec: Path-Aware Specification Generation via Large Language Models

This work introduces Path2Spec, a divide-and-conquer framework that leverages LLMs to extract all execution paths from an input program, generates path-specific specifications for each, and merges them into a comprehensive overall specification.

Dan Huang, Zhensu Sun, Hui-Hui Huang et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.