Skip to content

Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems

Aug 2026 · 0 citations
Computer Science

TL;DR

This work defines a grouping metric, specify a harness, and shows how tracking a human-AI pair's grouping over time yields the compounding signal that Paper 1's field study requires.

Abstract

Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue this measures the wrong axis. The models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision: how tightly concentrated their outputs are around that target across repeated, identical requests. Borrowing the marksman's distinction, capability is where the average shot lands; reliability is the size of the group. I make three claims. First, precision, not capability, is the frontier differentiator between systems, and benchmark culture systematically fails to measure it, reporting central tendency rather than spread. Second, precision is measurable, cheaply and without circularity, by running a fixed suite of deterministically scored tasks many times at fixed temperature and computing the per-task consistency of outcomes -- no model-in-the-loop grader required. Third, the measurement is not merely descriptive but decision-guiding: it separates consistent failures (a tight group off-centre, correctable by the operating discipline of Paper 1 -- a sight adjustment) from scattered failures (a wide group, correctable only by changing the model or its sampling -- a rifle problem). I define a grouping metric, specify a harness, and show how tracking a human-AI pair's grouping over time yields the compounding signal that Paper 1's field study requires. A first real run, since replicated, illustrates both the method and its most important limit: one measured gap was closed completely by a single rule (0/5 ->5/5), while a suite of tasks authored from the rules themselves found no value, because a frontier model already embodies explicit good practice -- establishing that a discipline's worth is found by measurement on real work, not constructed from its own rulebook.

View source

Similar papers

Preprint Jul 2026

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment

The experimental record on human-AI collaboration shows that naive combination often underperforms the stronger partner, implying that the human contribution must be repositioned toward specification, verification, and oversight, a shift visible in experiments but, so far, barely visible in field labor-market data.

Ancuta Margondai, Julie Rader, E. Rader et al. · 0 citations
Preprint Aug 2026

Split the Labor: Separating Evidence Interpretation from Decision Aggregation

This work proposes a four-field evidence tuple (hypothesis, reliability bucket, rationale, provenance, provenance) and shows that fixing it determines both halves of the framework and applies to a class of rules beyond language models: score-summing triage engines, diagnostic panels scored by counting positives, and additive multi-signal detectors.

Zhelun Wu · 0 citations
Preprint Jul 2026

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. We find that most of the apparent shift in gains toward harder tasks does not reflect a change in the shape of the difficulty-response curve. On METR time-horizon data, a single Rasch model with rising ability reproduces this pattern, so it is largely explained by ceiling effects rather than a qualitative change in capability. This echoes how the choice of metric can make claimed emergent abilities look like a property of the models themselves. We then identify a smaller hard-task effect that survives this control. Isolating it is difficult on agentic benchmarks, because newer models are usually run with newer agentic harnesses, so a gain on hard tasks cannot be assigned to the model or its scaffold. We break the confound with LiveCodeBench, a public competitive programming benchmark that runs no agentic scaffold while pairing dated models with an exogenous difficulty ordering. After accounting for the rise in overall ability, models released after September 2024 still gain on the hardest problems beyond what their easy and medium performance predicts, by about +0.40 logits under our most conservative assumption, raising the hard-problem solve rate from roughly 18% to 25%. The effect is led by the strongest reasoning models and holds for hard tasks that need only short reasoning, not autonomy over long horizons. We present this as a result specific to competitive programming, since our clean identification rests on a single coding benchmark. We release the LiveCodeBench Difficulty Panel (66 dated models x 1,055 problems) and our analysis code.

Hanwen Xing, Pengyu Wang, Bingxu Meng et al. · 1 citation
Preprint Aug 2026

Certifying Compressed Language Models: An Audit and a Statistical Toolkit

Paired equivalence testing at a declared margin is supply: paired equivalence testing at a declared margin, with certification tables giving the items an evaluation needs, computed from disagreement observed under compression, not from independent-binomial variance.

Amogh Singh · 0 citations
Review

Generative AI and the Global Redistribution of Expertise

A task-level expertise database for ISCO-08, the international standard that allows for cross-country comparisons, is built and shows that generative AI reaches the tasks that make an occupation expert in some occupational groups but not others.

Paweł Gmyrek, Héctor Segura, Hernán Winkler et al. · 1 citation
Preprint Jul 2026

Optimization Is Not All You Need

In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text to aid the detection of machine-generated text.

Minh Hua, Rita Raley · 0 citations

Related blog posts