SemanticAlign-Bench (SA-Bench), a diagnostic benchmark covering 30 papers from ICLR, ICML and NeurIPS 2025, is introduced and indicates that scaffolds optimized for executability provide limited leverage for scientific reproduction; narrowing the gap requires scaffolds that prioritize semantic specification verification.
Abstract
LLM agents can generate paper reproduction code, yet often produce scientifically unfaithful implementations. We define this failure mode as semantic drift, where generated code silently diverges from the paper's specifications. We introduce SemanticAlign-Bench(SA-Bench), a diagnostic benchmark covering 30 papers from ICLR, ICML and NeurIPS 2025. For each paper, we decompose its specifications into atomic and verifiable implementation claims, which we call Semantic Alignment Units (SAUs) and evaluate repositories along four diagnostic dimensions spanning numerical, methodological, protocol and ordering drift. In total, we construct 1,491 SAUs across five ML domains and evaluate 12 generator configurations (4 models $\times$ 3 scaffolds). Even the strongest configuration (Claude+PaperCoder) achieves a mean SAU score of only 0.301 out of 1.0, with an overall mean of 0.221 across 360 evaluations. A failure taxonomy reveals that agents attempt most requirements but implement them incorrectly, with implementation mismatch and stubs accounting for the majority of zero-scored claims. Our analysis further indicates that scaffolds optimized for executability provide limited leverage for scientific reproduction; narrowing the gap requires scaffolds that prioritize semantic specification verification. The benchmark, annotations and evaluation pipeline are publicly available.
SemVerBench is introduced, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo), and six frontier models are evaluated: Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar).
Large Language Models (LLMs) have shown promising performance in generating Object Constraint Language (OCL) constraints from natural language specifications. However, existing evaluations rely on publicly available UML models, which may overestimate generalization due to potential data leakage and reliance on recurrin...
Hamza Attarwala, Moataz Chouchen, Omar Alam et al.· Proceedings of the ACM/IEEE...· 0 citations
Aggregate factuality scores hide where a language model succeeds, which relations it confuses, and whether an answer survives innocuous changes to the question or decoder. We introduce PROOF, a profile-oriented benchmark for factual coverage in instruction-tuned language models. PROOF converts a frozen Wikidata snapsho...
A. Chetvergov, Mikhail Solovev, Timofei Sivoraksha et al.· 0 citations
A structurally valid DeFi workflow can still authorize a costly trade. We introduce DeFiFlowBench, a benchmark of 207 team-authored prompts for natural-language DeFi workflow synthesis. It measures graph coverage, configuration completeness, and declared safety predicates, then tests supported trade configurations on a...
Abhinav Kumar, Harshit Arora, Varun Singh et al.· 0 citations
TeXFix-Bench is presented, a multi-format benchmark for LLM-based full-source document repair grounded in a mined fault taxonomy and releases the taxonomy, DocMut, and all campaign artifacts.
The results show that point accuracy alone is insufficient for characterizing LLM reliability in assertion generation and motivate robustness-aware evaluation for AI-assisted hardware verification.
Fnu Aditi· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.