The ANVIL compiler architecture is proposed, where ANVIL performs exceptionally well on leading benchmarks by using language models purely where they are effective, rather than as a catch-all tool.
Abstract
Systems that turn natural-language descriptions of optimization problems into solver-ready code generally use a language model at every stage, including the final translation from a mathematical formulation into executable model-building code. We propose the ANVIL compiler architecture, where we separate these concerns. A language model is called only once, assisted by constraint guidance based on problem type, to produce a LaTeX formulation. A deterministic compiler then translates that LaTeX into code with no language model involvement. We describe the deterministic compiler (a normalizer, a recursive-descent parser producing a typed intermediate representation, analysis passes that bind symbols to a dataset schema, and a code emitter) and evaluate it on the 354 easy and hard problems of the NLP4LP benchmark. The compiler produced code for 329 of 354 formulations (92.9%), taking the deterministic path in every one of those cases and never falling back to model-generated code. Median compile time was below the 10ms resolution of our timer. Overall, our formulations achieved an accuracy of 98.9% over easy problems and 91.1% for hard problems. The gap between these compilation and accuracy figures is a key point of analysis, and we analyze it: formulations that failed to compile, problems that returned as infeasible, problems raising errors at runtime, and problems returning a wrong objective. Almost all of these errors trace back to the formulation rather than to the translation. ANVIL performs exceptionally well on leading benchmarks by using language models purely where they are effective, rather than as a catch-all tool.
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At c...
Yun-Tian Deng, Peng-Yu Nie, Stuart M. Shieber· 0 citations
This work is a continuation, and generalization, of Andrew Appel's landmark work on “Compiling with Continuations”, but instead of natural-deduction-based languages like the lambda calculus, it uses sequent-calculus-inspired languages throughout all intermediate stages.
Marius Müller, David Binder, Marco Tzschentke et al.· ACM Transactions on Programm...· 0 citations
Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question...
Yuan Li, Han-Yun Jiang, Guo-Wei Tian et al.· 0 citations
Compilable Academic Document Parsing (CADP) is proposed, a paradigm that reconstructs a full page as contextual \LaTeX{} plus executable Python, so that structure-preserving elements and executable chart representations can be reconstructed, recompiled, and directly verified against the source page.
SDDL is introduced, a neuro-symbolic framework that translates natural-language scheduling problems into compact, solver-aligned representations of tasks, resources, constraints, and objectives, while delegating low-level modeling and search to a deterministic compiler and external solver.
FORM is a domain-specific symbolic manipulation language widely used in particle physics for processing the very large algebraic expressions arising from multi-loop Feynman diagram calculations. Despite its central role in precision theoretical physics, no artificial-intelligence tooling exists, to our knowledge, for a...
B. Chargeishvili· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.