Skill-based Agentic Evaluation for Real-time Data Science Tasks
A framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring, which achieves a 29% improvement in the Matthews Correlation Coefficient and a 16% reduction in token consumption per test case.
Aniruddha Tamhane, Raghavendra Addanki, Ayushi Aggarwal et al.
· 0 citations