Skip to content
Preprint

FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training

Aug 2026 · 0 citations · 26 references
Computer Science

TL;DR

FlavourBench is introduced, which instead compiles dense answer maps from a versioned culinary environment and remains similar under alternative metrics, task filters, family weights, and three public Epicure checkpoints.

Abstract

Open-ended language-model evaluation often substitutes another model or a small preference panel for a missing answer key. We introduce FlavourBench, which instead compiles dense answer maps from a versioned culinary environment. Each task asks for a three-ingredient portfolio from eight candidates; before inference, Epicure scores all 56 portfolios. We evaluate 27 frontier endpoints on the same 534 substitution, pairing, and constraint tasks, yielding 14,418 complete model-task observations. Anchor-cluster bootstraps and multiplicity-controlled paired tests resolve 101 of 351 model contrasts. Grok 4.6 has the largest point estimate at 65.1, but the corrected evidence does not identify a unique best endpoint. The ranking replicates across independently compiled panels and remains similar under alternative metrics, task filters, family weights, and three public Epicure checkpoints. We then run a preregistered, three-seed post-training study. LoRA SFT of a pinned Qwen3-0.6B checkpoint on 270 Epicure-optimal answers improves its score on 84 anchor-disjoint maps by 13.30 points over a format- and label-matched control (95% CI 6.52 to 20.29, p = 0.000170), and the effect replicates on all 534 public maps. The replication gain is 11.73 points (95% CI 8.98 to 14.54). On the primary split, both trained arms parse every response while the format control does not improve on the base model. The release contains prompts, exhaustive reward maps, raw responses, training and evaluation manifests, statistical plans, code, and an offline verifier.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck

The first compute-normalised comparison of five TTS families across five open-ended generation benchmarks spanning medicine, law, finance, general chat, and creative writing is conducted - grounded in a unified framework that decomposes the effectiveness of each method's token budget into exploration and exploitation.

Davide Romano, Kanak Raj, Jerrod Parker et al. · 0 citations
Preprint Jul 2026

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

Under cross-benchmark transfer evaluation across twelve models completed by all four readouts, difference-in-means attains the highest model-level mean on ten models, while logistic regression wins the most capability-model cells.

Yanshi Li, Xue Bai, Shuman Liu et al. · 0 citations
Review Jul 2026

ConfidenceBench: Evaluating Confidence Calibration in Large Language Models

ConfidenceBench, a calibration benchmark that evaluates verbalized confidence estimates in 15 frontier LLMs using the Brier score, a proper scoring rule that incentivises truthful probability reporting shows that verbalized confidence calibration is a distinct and practically important axis of LLM reliability, complementary to standard accuracy-based evaluation.

M. ffrench-Constant, Daniel Yang, Xinmeng Huang et al. · 1 citation · ⚡1
Preprint Aug 2026

UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

Organizations maintain task-specific adapters for open-weight language models, and each new base-model release forces a migration decision: retain existing specialists, port adapters, refresh from preserved behavior, or retrain. Prior transfer work evaluates isolated model pairs, without studying these choices across real model release sequences. We present UpgradeBench, a decision-driven longitudinal benchmark covering four consecutive Qwen releases, one continuation checkpoint, six tasks, and two model sizes, augmented by OLMo checkpoints with known training lineage. The benchmark disentangles three core questions: whether a new checkpoint improves fixed-recipe retrained specialist performance, whether specialization assets transfer across versions, and what recovery resources are usable. We observe upgrade gains differ across task-scale-release episodes: some retrained baselines improve while others stay within training noise, with durability ranging from under one release interval for text-to-SQL to over fourteen months for intent classification. Direct adapter copying depends neither on architecture nor model family: on OLMo, retention drops from 0.88-0.99 at 46B-token continued pretraining to zero at 2.9T tokens; annealing and model souping introduce no extra harm, with portability decaying with continued-pretraining distance. Given preserved input data, teacher relabeling recovers target-base specialists without fresh gold annotations, though compute savings are not guaranteed. Simulating a fixed decision policy over 33 upgrade episodes yields 0.37pp mean quality regret with zero behavioral regressions at one-third the compute and label cost of full retraining. A lightweight CKA probe over 256 prompts predicts cross-version adapter portability (Spearman 0.74 across eight model pairs). We release per-example predictions, cost logs, split manifests, and evaluation code.

Yefei Chen, Wei-Ning Zhang · 0 citations