Skip to content

IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

Sep 2026 · 0 citations · 15 references
Computer Science

TL;DR

A bounded IHEval comparison uses the same SmolLM2 checkpoint and output budget while preserving its published instruction roles and scorer, and the benchmark measures conditional task and output-contract success.

Abstract

An external record may contain a procedure to apply or text to read, depending on the user's request. IBBench-Light tests both uses against the same record. Twelve semantic bases yield 144 matched pairs per model; four quantized instruction models produced 1,152 archived greedy responses. Paired exact-contract accuracy (PECA) requires both members to satisfy their output contracts. Qwen succeeds on 132 execute and 109 process prompts, but only 97 complete pairs, showing what marginal averages omit. We audit literal-target exposure and case normalization, then add 1,722 logged CPU generations to test directive-absent controls, twelve additional semantic bases, within-base wording changes, and generation stopping. In the pinned Phi rerun, changing the end-of-sequence (EOS) set changes exact paired success from 0/144 to 62/144. A bounded IHEval comparison uses the same SmolLM2 checkpoint and output budget while preserving its published instruction roles and scorer. The benchmark measures conditional task and output-contract success. Its task margins and paired count need to be read together with the stopping policy.

View source

Similar papers

Preprint Aug 2026

QuoteBench: How Matched Scores Can Hide Command-Path Failures

Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.

Shangao Li, Yao Zhang, Volker Tresp et al. · 1 citation · ⚡1
#artificial intelligence Preprint Sep 2026

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

A memory can answer a current query correctly while discarding distinctions required by a later update. We investigate this failure with a paired-history audit: two histories have the same current answer, receive a shared future update, and require different subsequent answers. A pilot evaluates 24 history pairs across...

Guang-Zhe Zhang · 0 citations
#artificial intelligence Preprint Sep 2026

Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory

A model that inherits one-line memories may pull one archived source record before acting; a directive in the store can steer that pull: a pointer, a criterion or both. Across sixteen registered studies (179,352 attempts) we measured where the request goes under each form; every result is descriptive, with registered i...

Kazuki Nakayashiki · 0 citations
#artificial intelligence Preprint Sep 2026

Adaptive-GEPA: Make Your Harness Fit Heterogeneous Requests

Reflective optimizers such as GEPA improve language model prompts from execution traces and evaluator feedback; full-program extensions can also rewrite tools and control flow. In practice, a user hands the same endpoint heterogeneous requests whose effective solutions require different tools, reasoning modes, and cont...

Tian-Yu Chen, Yasi Zhang, Rui-Yi Wang et al. · 0 citations
#artificial intelligence Review Oct 2026

Language-Conditioned Token and Reasoning Efficiency in Large Language Models: A Paired Cross-Lingual Study Protocol

Large language models incur language-dependent representation and inference costs, but existing comparisons often conflate input language, assigned observable-trace language, and answer realization. We specify a prospective paired study that separates these interfaces while holding the semantic item, checkpoint, and an...

Gen-Liang Zhu, Chu Wang · 0 citations
#artificial intelligence Preprint Sep 2026

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

SemVerBench is introduced, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo), and six frontier models are evaluated: Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar).

Qi-Bai Chen, Ze-Ming Liu · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.