AlGOBENCH is introduced, a framework that automatically builds novel algorithmic problems from known competitive-programming problems through structured constraint-shifting transformations and error analysis shows that failures are mainly algorithmic rather than implementation-level, suggesting that ALGOBENCH evaluates adaptation beyond functional correctness.
WM-SAR, a spectral subgraph repair method that estimates node-edge amplification, greedily grows a connected repair region by marginal residual-spectral relief, and sends only this region to an LLM for root-cause repair, achieves stronger long-horizon stabilization and root-cause recovery under compact token budgets.
Grounded Iterative Language Planning (GILP), which trains only a small parameterized backbone and combines it with API-based agent reasoning, and a consistency gate asks for revision when the two disagree are compared.
Experiments on competitive programming and combinatorial optimization benchmarks show that AlgoSkill improves over direct LLM generation, chain-of-thought prompting, self-refinement, and MCTS without typed skills, which support treating automatic algorithm design as verification-guided skill scheduling rather than one-shot code generation.
Xinyuan Song, Z. Cai, Liang Zhao· arXiv.org· 1 citation