Skip to content
Preprint

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment

Jul 2026 · 0 citations · 14 references
Computer Science

TL;DR

The experimental record on human-AI collaboration shows that naive combination often underperforms the stronger partner, implying that the human contribution must be repositioned toward specification, verification, and oversight, a shift visible in experiments but, so far, barely visible in field labor-market data.

Abstract

Between 2023 and 2026, frontier AI systems crossed documented human expert baselines on a growing set of bounded, well-specified, evaluable cognitive tasks, including graduate-level science questions, competition mathematics, software-engineering benchmarks, and structured diagnostic reasoning, while the length of tasks such systems can complete at 50% reliability doubled roughly every seven months. These crossings are rapid and broad, but the frontier is jagged: humans retain decisive advantages in long-horizon reliability, genuinely novel problems, calibrated self-knowledge, sample-efficient learning, and embodied action, and benchmark results overstate deployed capability for reasons that are themselves now documented, namely contamination, construct validity, vendor self-evaluation, and the gap between 50% reliability and the reliability that economic work requires. Concurrently, humans increasingly use these systems as cognitive extensions. The offloading literature predicts costs to unaided skill, and early field evidence is consistent with such costs, though the largest meta-analytic evidence on prior technologies points the other way, and the question of whether generative AI differs is open. Finally, the experimental record on human-AI collaboration shows that naive combination often underperforms the stronger partner, implying that the human contribution must be repositioned toward specification, verification, and oversight, a shift visible in experiments but, so far, barely visible in field labor-market data. This paper states the resulting position, rapid crossings on a jagged frontier with a human role that must be redesigned rather than defended, and draws out its theoretical and practical implications.

View source

Similar papers

Preprint Jul 2026

What AI Red-Team Evaluations Can and Cannot Prove

This work defines the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and uses it to locate that boundary exactly.

Bandana Kaur · 0 citations
Book

Meta-Cognitive Judgment

Unknown authors · 0 citations
Preprint Aug 2026

AI and the Research Team

Artificial intelligence is associated with larger research teams, yet in mathematics, among the most codifiable fields, individual researchers working with AI now produce research-grade results. A span-of-control model reconciles these observations. AI lowers execution cost, which expands laboratory scale, and automates codifiable tasks, which lowers the member share of each unit. Team size is therefore quasi-concave in AI capability, with at most one peak. The model predicts that a fully codifiable team peaks when effective automation coverage reaches a closed-form threshold, typically near complete coverage, and, among fields with shared primitives that possess an interior peak, those with less irreducibly human task content peak first. Under explicit priors, the 90 percent forecast intervals for the fully codifiable peak span 2026 to 2030.

Johan Fourie · 1 citation
Review

Generative AI and the Global Redistribution of Expertise

A task-level expertise database for ISCO-08, the international standard that allows for cross-country comparisons, is built and shows that generative AI reaches the tasks that make an occupation expert in some occupational groups but not others.

Paweł Gmyrek, Héctor Segura, Hernán Winkler et al. · 1 citation
Review Open access Aug 2026

A Review of Human-AI Complementarities Across Multiple Dimensions of Organisational Complexity

A five-dimensional diagnostic framework that maps the challenges of human-AI collaboration across Integration, Representation, Scale, Temporality, and Adequacy gaps and shows that augmentation remains the dominant and most viable mode of use in complex environments.

Ganesh Sankaran, Marco A. Palomino, G. Siestrup · 0 citations