Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 13423-13424· 0 citations· 22 references
TL;DR
This workshop seeks to consolidate efforts by providing an interdisciplinary forum for presenting cutting-edge research, sharing deployment experiences, and showcasing real-world systems in this rapidly evolving field of AI Data Scientist.
Abstract
As data volumes and analytical demands grow, traditional data science workflows struggle to meet the need for efficiency, scalability, and reliability. The rapid advancement of large language models (LLMs) has opened new possibilities for AI-powered agents to augment or automate end-to-end data science pipelines—from data exploration and cleaning to modeling, evaluation, and deployment. This emerging paradigm, termed the AI Data Scientist, has gained significant attention in research and industry, yet discussions remain fragmented regarding its integration, evaluation, and real-world impact. This workshop seeks to consolidate these efforts by providing an interdisciplinary forum for presenting cutting-edge research, sharing deployment experiences, and showcasing real-world systems. The workshop will feature invited talks, paper presentations, a demo track, and a panel discussion, aiming to foster community-building and guide responsible development in this rapidly evolving field.
This paper examines Data Agents from a harness-centric perspective, introducing a taxonomy of Data Agents and associated data environments, and identifying four open reliability problems: inactive semantic calibration, missing clarification, missing experience transfer, and the missing verification-repair repository.
Hua-Chi Zhou, Yu-Jing Zhang, Jia-He Du et al.· 0 citations
Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows. However, using proprietary commercial AI systems raises concerns about transparency, reproducibility and privacy, which are essential for scientific practices. To this e...
J. Stark, S. Saikrishnan, Vikram Seenivasan et al.· 1 citation
A large-scale benchmark with 60\sim 90× more queries than prior work, built on 3× more databases, an automated pipeline that can execute existing methods without manual intervention, and multi-dimensional, fine-grained evaluation metrics for comprehensive assessment.
Bo Li, Chenzhan Wang, Long-Kang Lin et al.· Proceedings of the 32nd ACM...· 0 citations
Modern organizations depend on data platforms that must serve two audiences at once: analytical workloads that expect consistent, well-governed tables, and artificial intelligence (AI) workloads that expect fresh, versioned, feature-ready data delivered to training and inference systems. Most platforms were not designe...
Santoshi Maddali· World Journal of Advanced Re...· 0 citations
An overview of AI’s role across the stages of scientific research is provided, illustrating how AI can assist researchers by aggregating and synthesizing a large body of work across various domains, supporting methodological implementation, and facilitating communication and publication.
Neda Sadeghi, Erin Nakamura, Luke J. Norman et al.· Aperture Neuro· 0 citations
This work shows how machine learning and GenAI can be used to assist with two specific tasks: First, when reading CSV files, it needs to be decided whether the first row is a header or not, and how machine learning and GenAI can be used to assist with two specific tasks.
Alexander van Renen, Moritz Rengert, Macallyster Edmondson et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.