Skip to content
Preprint

TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair

Aug 2026 · 0 citations · 22 references
Computer Science

TL;DR

TeXFix-Bench is presented, a multi-format benchmark for LLM-based full-source document repair grounded in a mined fault taxonomy and releases the taxonomy, DocMut, and all campaign artifacts.

Abstract

Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatched environments, broken imports, or package conflicts. Existing document-repair evaluations inject faults with ad-hoc edits that lack an empirical fault model. We present TeXFix-Bench, a multi-format benchmark for LLM-based full-source document repair grounded in a mined fault taxonomy. A Grounded-Theory study of localized hard-crash LaTeX faults from TeX Stack Exchange, GitHub commits, and package documentation (168 verified faults, dual open coding at $\kappa$=0.34) yields an 18-category taxonomy instantiated as DocMut: 48 AST-aware operators across three formats. A three-model cross-benchmark shows DocMut faults are 5.6-9.2 pp harder to repair than pattern-based mutations on the same seeds, and a real-error case study (88 mined human crashes, 67.0% repair success) brackets both synthetic sets from below. We construct 10,437 instances from 743 openly licensed seeds and evaluate seven LLMs under a fixed zero-shot protocol with provider-pinned routing, collecting 48,651 attempts at about USD 200 total inference cost. A complete 6,613-instance x 7-model balanced matrix confirms all rankings. A pinned engine gate yields a 27.5-point intention-to-treat compile spread (56.7-84.2%). Typst is markedly harder than LaTeX and Markdown. A restoration oracle over 28,129 compiling repairs shows that 13.6-18.5% of compiling repairs materially alter document text, and restoration rank diverges from compile rank: the model with the lowest compile rate restores content best among its successes. Compile success alone overstates repair quality. We release the taxonomy, DocMut, and all campaign artifacts.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Externalizing Requirement-to-Repair Artifacts as Observable Traces for LLM-Based Program Repair

Repository-level repair requires not only correct patches but also inspectable records that explain how issue requirements are translated into code changes and post-edit evidence. We contribute THEMIS, a stage-aware repair workflow that externalizes this requirement-to-repair process through semantic interpretation, a...

Ze-Wen Tao, Shin-nosuke Ishikawa · 0 citations
#artificial intelligence Preprint Sep 2026

Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

Large language models (LLMs) are increasingly applied to the automated repair of C/C++ security vulnerabilities, and compile rate is a commonly reported proxy for progress: whether the generated patch compiles. We argue that compile rate is a scientifically unreliable metric for single-function vulnerability repair, an...

Om Nepal, Sushant Aryal, Oluseyi Olukola et al. · 0 citations
Preprint Aug 2026

PatchWrite: One Line, Not One Section -- Compile-Gated, Validity-Preserving Editing for AI-Drafted Manuscripts

PatchWrite constrains how candidate edits become committed manuscript states: it reuses bounded EDIT N M editing and rollback, but tightens compilation acceptance with fatal-log checks and adds evidence locks that require every cited key and experimental numeric token to be attested by a reference registry or experimen...

Wei-Wei Yang · 0 citations
Open access 2026

Deployment-Oriented Evaluation of LLM-Generated Optimization Code: Repairability, Compatibility, and Constraint-Coverage in Parallel-Machine Scheduling

Large language models (LLMs) increasingly generate executable optimization code, yet evaluations often rank programs by objective value, overlooking deployment-relevant properties such as validity under structural constraint changes, failure localization, repairability, and component compatibility. We introduce GenSE-S...

Achraf Ghorbel, Nourchène Elleuch Ben Ayed, Keletso J. Letsholo et al. · 0 citations
Review Aug 2026

OdinEval: A Reproducible Benchmark for LLM-Based Program Repair in the Odin Programming Language

OdinEval is presented, a reproducible benchmark built from documented defects in public Odin repositories built from documented defects in public Odin repositories, that evaluates six language models on 168 filtered instances under one shared protocol.

Bang Xie, Hao Liu, Zhi-Yuan Peng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.