Preprint
Jul 2026
TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA Specification Generation
The main finding is about measurement itself: an exact oracle gives not one correctness number but a range, which is called the correctness envelope, and the findings inside it are stable.
Arslan Bisharat, E. Spencer, B. Ortiz et al.
· 0 citations