The introduction of ScientistTwo, a fully autonomous multi-agent framework designed to realize problem-driven autonomous research, and its results show that ScientistTwo is not merely an assistive tool but an autonomous scientific pioneer capable of pushing the frontiers of human discovery.
Abstract
Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory. The ultimate vision for AI in science is problem-driven autonomous research: given a fundamental challenge by a human expert, the AI independently navigates the scientific landscape, uncovers theoretical and empirical bottlenecks, and systematically expands the frontier of knowledge. In this paper, we introduce ScientistTwo, a fully autonomous multi-agent framework designed to realize this vision. Specifically, ScientistTwo takes an initial problem as input, establishes state-of-the-art baselines, formulates novel hypotheses, and coordinates specialized agents to orchestrate an end-to-end discovery cycle without human intervention. Moreover, the framework rigorously conducts experiments using diverse datasets and metrics, refines methodologies through automated ablation studies, and validates research findings via a closed-loop simulated peer-review rebuttal engine. To evaluate ScientistTwo’s capabilities against the highest standards of human scientific achievement, we benchmark it across papers accepted at top-tier conferences such as ICLR, ICML, and NeurIPS. As a result, ScientistTwo autonomously generates expert-level, publishable papers and fully verified, executable codebases. Its solutions consistently outperform human state-of-the-art models, and achieve higher average review ratings than human-authored papers under automated AI review agents. These results show that ScientistTwo is not merely an assistive tool but an autonomous scientific pioneer capable of pushing the frontiers of human discovery.
ABSTRACT Artificial intelligence systems are transforming scientific discovery by accelerating specific research tasks, from protein structure prediction to materials design, yet remain confined to narrow domains requiring substantial human oversight. Exponential growth of scientific literature and increasing domain sp...
Gabrielle Wehr, Reuben Rideaux, Amaya J. Fox et al.· Advancement of science· 0 citations
This paper explores the emergence of artificial research intelligence (ARI), autonomous AI systems designed to conduct end-to-end scientific research, from ideation to manuscript production. Unlike simple AI enhancements, ARI utilizes ensembles of agentic models to automate the scientific method, potentially creating a...
Michael Ridley· Information Technology and L...· 0 citations
Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely assess it by asking models to generate ideas from a static, curated set of reference papers. That passive setup departs from the retrieval-and-...
Yunxiang Mo, Tianshi ZHENG, Yi-Sen Gao et al.· 0 citations
Mechanist is an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence, and develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretr...
Mengru Wang, Jun-Feng Fang, Shuo-Fei Qiao et al.· 1 citation
This survey bridges the existing gap by presenting a comprehensive blueprint for scientific agents' design and introduces a unified taxonomy based on capability envelope and capability maturity, characterizing both the scope of scientific workflow coverage and the reliability of agent behavior under realistic research...
Xin-Ming Wang, Jian Xu, Sheng Lian et al.· IEEE Transactions on Pattern...· 11 citations
The Little Scientist is presented, a framework in which a LLM agent stepping through the scientific method can discover both new algorithms and new ensemble strategies that outperform prior solutions, and is demonstrated on two problems that require fundamentally different modes of discovery.
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.