Skip to content

Skill-based Agentic Evaluation for Real-time Data Science Tasks

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

A framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring, which achieves a 29% improvement in the Matthews Correlation Coefficient and a 16% reduction in token consumption per test case.

Abstract

We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring. Consider this example query:"what were last week's audience sizes"---the reference answer changes as the underlying data changes, so static references become outdated and standard LLM-as-a-judge pipelines cannot verify responses against a fixed ground truth. Our central contribution, ground-truth-as-code, encodes each expected answer as an executable reference function that recomputes the answer directly from live data at evaluation time, ensuring the reference remains consistent with the system it describes. We combine this with a factoid-level, format-agnostic judge that decomposes both the agent's response and the computed ground truth into atomic claims and scores precision, recall, and accuracy over them, irrespective of the response format (prose, list, table, HTML, etc.). The approach is applicable to agents whose expected outputs can be expressed as executable data computations. We validate the framework through a human--LLM agreement study on an internally developed machine learning skill deployed in production, using a synthetic database constructed to reproduce production schemas and entity relationships. Relative to a natural-language ground-truth baseline, our method achieves a 29% improvement in the Matthews Correlation Coefficient (MCC)---a class-balanced measure of agreement between expert annotators and LLM-as-a-judge predictions---and a 16% reduction in token consumption per test case, while a self-directed baseline lacking explicit ground truth is anti-correlated with human judgment. Agents that perform multi-source data integration and computation over non-stationary data are routinely deployed in industry; we propose ground-truth-as-code as a practical methodology for their evaluation.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

DolphinBench: Mapping the Pareto Frontier of Agent Memory

DolphinBench is presented, a benchmark that evaluates memory directly through an agent's task completion and requires all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically.

Soumil Rathi, Deshraj Yadav, Taranjeet Singh · 0 citations

Toward Self-Evolving Data Agents for Autonomous Data Analysis

Comparisons against stronger model and coding-agent competitors further indicate that both domain-specific agent runtime structure and foundation-model strength matter for autonomous data analysis.

Junhao Zhu, Lu Chen · 0 citations
Preprint Aug 2026

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations natu...

Weiliang Chen, Haowen Sun, Jun Gao et al. · 2 citations
#artificial intelligence Preprint Sep 2026

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are often tied to specific task types, execution environ...

Yu Liu, Zhi-Lin Liu, Zhi-Wei Yang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

K-Bench: measuring model performance on real scientific agent requests

K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs is reported, arguing that the informative quantity for scientific agents is not a leaderboard position but the joi...

Aubrey M. Brueckner, Darshil Patel, Yu-Huan He et al. · 2 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.