Skip to content

OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets?

Sep 2026 · 0 citations · 33 references
Computer Science

TL;DR

OptiArena is introduced, a budget-controlled testbed for studying whether LLMs can improve executable game-playing algorithms through five rounds of code edits within a fixed minimal scaffold and under bounded evaluator feedback and fixed resource budgets.

Abstract

Static QA and code-generation benchmarks only partially capture the role that large language models (LLMs) now play as coding agents and research tools. We introduce OptiArena, a budget-controlled testbed for studying whether LLMs can improve executable game-playing algorithms through five rounds of code edits within a fixed minimal scaffold and under bounded evaluator feedback and fixed resource budgets. The testbed uses two optimization regimes, surface obfuscation controls, calibrated references, held-out/stress splits, and diagnostics for degradation and exceptional failures, with LLM API cost reported separately from local evaluator wall-clock. The empirical study asks three questions: whether models can close the calibrated gap between a designated weak starter and an editable competent baseline, whether they can refine editable competent baselines without damaging them, and whether gains survive surface obfuscation controls. Across twelve frontier LLMs and five games, models improve designated weak starters more consistently than they refine editable competent baselines, with substantial variation across games and models. OptiArena provides a practical testbed for measuring bounded-resource algorithm optimization within the five-edit, fixed-scaffold setting studied here. Code is available at https://github.com/WJ-Peng/OptiArena.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Benchmarking Prompt Optimization of Large Language Models With Chess

A chess benchmark built from 1,118 Lichess puzzles is introduced to study APO for frozen LLMs: it is used to evaluate six APO algorithms on eight target models, measuring not only baseline strength but also how much each model responds to optimization and whether optimized prompts transfer across models and to game pla...

Timothée Lesort, Alejandra López de Aberasturi-Gómez, Tristan Karch et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SAILOR: Solver-Assisted Interactive LLM-based Optimization Recovery

Natural-language descriptions of optimization problems may be incomplete or vague about numerical information that a solver requires, including costs, capacities, demands, bounds, and penalties. A language model can translate the description into code, but when a required value is absent it must either stop or guess. W...

Shaghayegh Sadeghi, Steve Smith, D. C. Del Rey Fernández · 0 citations
#artificial intelligence Preprint Sep 2026

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.

Chao Wang · 0 citations
Review Sep 2026

ExecuCritic: Calibrated Critic Shaping for Code Generation with Verifiable Rewards

Execution feedback is a useful supervision signal for code models because unit tests are objective and directly measure program correctness. Its weakness is that an entire program is often reduced to one pass or fail bit, leaving RLVR to solve a difficult credit assignment problem. At the same time, coding systems ofte...

Jun Cao, Ying-Jie He · 0 citations
#machine learning Preprint Sep 2026

How Reusable Are Benchmarks with Richer Feedback?

We study whether benchmarks reliably guide model selection as developers adapt to evaluation feedback across multiple criteria. We find that the worst-case test-set size needed to estimate the best score among $k$ adaptively chosen models, under any convex combination of the criteria, grows exponentially with the numbe...

Youssef Allouah, John C. Duchi · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.