Benchmarking Prompt Optimization of Large Language Models With Chess
A chess benchmark built from 1,118 Lichess puzzles is introduced to study APO for frozen LLMs: it is used to evaluate six APO algorithms on eight target models, measuring not only baseline strength but also how much each model responds to optimization and whether optimized prompts transfer across models and to game pla...