It is argued that the key challenges are bridging the gap between personal and organizational memory, building effective agentic retrieval beyond naive RAG, and enabling organizations to evaluate their agents against their own data through self-serving evals.
This paper argues that genuine enterprise productivity gains require proactive agents: systems that surface relevant, actionable information to workers before they ask, and proposes the Context Graph, a live relational data structure that models enterprise entities, their relationships, and state transitions over time.
DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces, is introduced and key challenges for improving data-agent reliability are identified.
Boyan Li, Zhuowen Liang, Yupeng Xie et al.· 1 citation
Comparisons against stronger model and coding-agent competitors further indicate that both domain-specific agent runtime structure and foundation-model strength matter for autonomous data analysis.
DataClawEval is introduced, the first comprehensive benchmark designed specifically to evaluate the end-to-end task completion capabilities of autonomous agents in real-world data engineering scenarios, and it comprises 100 rigorous, end-to-end tasks spanning five execution engines.
Debin Meng, Jiaming Yang, Zefang Zong et al.· 0 citations
Experiments on public benchmarks and real-world industrial BI workloads show that QwenPaw-Data improves both verifiable data access capability and higher-level analytical quality, offering a practical foundation for reliable, traceable, and continuously improving enterprise data agents.
Tian Zeng, Yuntao Hong, Zhongjun Ding et al.· 1 citation
The AgenticData system, an agentic data system that enables natural-language query analytics over heterogeneous data sources, is introduced and its ability to handle diverse data sources accurately and efficiently is illustrated.
Peiyao Zhou, Ji Sun, Yaoqiang Xu et al.· 0 citations