Skip to content

AutoKD: Autonomous Knowledge Discovery

Sep 2026 · 0 citations · 35 references
Computer Science

TL;DR

This paper introduces AutoKD, a multi-agent framework for autonomous knowledge discovery that is both computational and cumulative, allowing validated findings to persist and inform subsequent inquiry.

Abstract

Scientific discovery in data-rich domains is currently constrained by human bandwidth: the growth in the volume and complexity of real-world data far outpaces the rate at which researchers can read, reason, and synthesize. Recent LLM-based multi-agent systems have begun to automate portions of the research cycle, but they target hypothesis generation in settings where validation cannot itself be automated, and each run is one-shot, with no mechanism for findings to accumulate or steer subsequent inquiry. This paper introduces AutoKD, a multi-agent framework for autonomous knowledge discovery that is both computational and cumulative, allowing validated findings to persist and inform subsequent inquiry. Six coordinated LLM agents collaborate in an open-ended discovery loop, where accepted findings are stored in a persistent insight graph that serves as both long-term memory and an exploration-steering mechanism. We evaluate AutoKD on three diverse datasets from two perspectives: Open-ended Quality against published findings, and Conditioned Quality via literature-derived queries. Across both evaluation perspectives, AutoKD covers known findings and surfaces substantive discoveries that complement human-driven research. Our code is available at https://github.com/GeQinwen/AutoKD.

View source

Similar papers

Preprint Aug 2026

ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond

Results on diverse long-horizon benchmarks demonstrate the efficacy of ScienceFlow's ability to sustain effective research processes, and demonstrates that efficient state management, adaptive exploration, and objective-aligned execution are critical for scaling autonomous research beyond short-horizon interactions.

Ming-Ming Zhao, Ji-Qian Dong, Kangping Xu et al. · 4 citations
Preprint Aug 2026

Multi-Agent Discovery and Resource-Aware Autonomous Exploration of Scientific Datasets

Modern scientific facilities and instruments generate datasets at scales that are difficult for individual researchers to discover, access, and explore. Although many datasets are publicly available, using them often requires familiarity with repository organization, data formats, multiresolution structures, and visual...

Aashish Panta, Hugo Lee, G. Scorzelli et al. · 0 citations
#artificial intelligence Review Open access Sep 2026

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

The introduction of ScientistTwo, a fully autonomous multi-agent framework designed to realize problem-driven autonomous research, and its results show that ScientistTwo is not merely an assistive tool but an autonomous scientific pioneer capable of pushing the frontiers of human discovery.

Jaehyun Nam, Jinsung Yoon, Yan Pan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human labour and on human-in-the-loop collabo...

Haotian Luo, Hao-Yu Wang, Ze-Yu Qin et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.