Skip to content
Open access

Scientific computing in the age of agentic AI: an exploratory field report

Aug 2026 · bioRxiv · 0 citations
Biology

TL;DR

Overall, it is found that the use of coding agents in scientific computing holds great promise for accelerating scientific research and increasing the reliability of critical systems, but that outstanding concerns remain.

Abstract

Scientific computing has become a central component of modern scientific discovery. Yet many computational tools are developed by small, specialized teams under incentives that encourage the release of rapidly prototyped tooling without commensurate attention to engineering concerns, including performance and maintainability. These gaps are particularly visible in the life sciences, where the advent of high-throughput sequencing and molecular profiling has made the production and processing of datasets routine at scales that strain reliability and cost. Recently, LLM-based agents have become increasingly capable, with publicly available systems possessing both significant domain knowledge in many scientific fields and the ability to autonomously operate over complex and specialized codebases in pursuit of well-defined goals. Together, these developments create a practical opportunity for scientific computing. Many of the persistent weaknesses of the scientific computing ecosystem stem from technical debt and a shortage of sustained engineering labor and expertise. Here, we examine coding agents as a potential way to address these weaknesses: we present an exploratory field report of eight early case studies in the application of LLM agents to scientific computing across a range of project scopes, from lightweight maintenance tasks to full performance-oriented rewrites of scientific libraries, with a focus on the life sciences. Each of these case studies is accompanied by reflections from the individual or group responsible for the work, including lessons from the process. Overall, we find that the use of coding agents in scientific computing holds great promise for accelerating scientific research and increasing the reliability of critical systems, but that outstanding concerns remain, including responsibility and ownership for such projects, and we suggest collaboration and stewardship with existing maintainers when feasible.

Read PDF

Similar papers

Review Open access Aug 2026

The Hitchhiker's Guide to Autonomous Research: A Survey of Scientific Agents.

This survey bridges the existing gap by presenting a comprehensive blueprint for scientific agents' design and introduces a unified taxonomy based on capability envelope and capability maturity, characterizing both the scope of scientific workflow coverage and the reliability of agent behavior under realistic research conditions.

Xinming Wang, Jian Xu, Sheng Lian et al. · 9 citations
Review Open access Aug 2026

A New Paradigm: Agentic AI for Scientific Discovery

This article examines the emerging paradigm of agentic AI for scientific discovery, traces the conceptual shift from tools to agents, lays out a six-stage workflow spanning literature synthesis to manuscript generation, and reviews practical systems in chemistry, equation discovery, materials science, and general machine learning research.

Alexander Taktakidze · 0 citations
Open access Nov 2025

Scientific discovery in the age of AI and supercomputing

Artificial intelligence (AI) and high-performance computing (HPC) are transforming scientific capabilities and the way science is conducted. Yet their combined impact on scientific discovery remains poorly understood, as do inequalities in access to these capabilities across countries and institutions. Drawing on metadata from more than five million scientific publications (2000–2024) across 27 fields, we examine how the convergence of AI and HPC correlates with scientific breakthroughs. Our results show that this computational synergy is most pronounced at the scientific frontier: research combining AI and HPC is more likely to introduce novel ideas and achieve top-cited status than either conventional work or research using AI or HPC in isolation. We also document growing disparities in access to supercomputing resources and AI expertise, which are increasingly concentrated in a small number of regions (dominated by the United States and China, though the EU27 aggregate maintains high competitiveness in combined AI+HPC output). The future of discovery will depend not only on advances in algorithms and computing power, but also on enacting policies that democratise these capabilities across the global scientific ecosystem.

Stefano Bianchini, A. Geuna, Fazliddin Shermatov · 1 citation
Review Open access Aug 2026

Quo Vadis? Scientific Discovery in the Age of Artificial Intelligence

The advancement of AI in science raises broader questions concerning the division of cognitive labour between human researchers and machines, and a typology of AI systems used in research is proposed, ranging from specialised scientific AI through scientific AI assistants and agents to hybrid experimental systems that combine computation and physical experimentation.

P. Jedlička · 0 citations
Preprint Aug 2026

BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks

Artificial intelligence (AI) promises to accelerate biological research by automating computational analyses. Yet the ability of AI agents to carry out computational biology at the scale of complete research studies has not been systematically evaluated. Here we introduce BixBench3, a benchmark that measures the capacity of AI agents to process raw biological data through to scientific results. We designed BixBench3 tasks to mirror the delegation of work from a scientist to an agent: the scientist chooses the research question and high-level methods, then delegates implementation of all analyses to the agent. In each task, an agent receives a research objective, methodological guidance, and raw data derived from a published scientific study, and must execute a sequence of analyses to achieve the research objective. The data artifacts resulting from these analyses - such as peak call matrices or differential expression tables - are programmatically graded against the corresponding artifacts generated and reported in the original study. Across 20 BixBench3 tasks encompassing the generation of 138 unique artifacts, we find that 13 frontier models achieve scores ranging from 0.00 for Gemini 3.1 Flash Lite to 0.48 for GPT 5.6 Sol. Agents perform worse on tasks with larger raw datasets (0.36 on tasks with<100 GB versus 0.10 on tasks with>100 GB) and on analyses requiring more sequential steps (0.36 at 1-2 steps vs 0.24 at 3+). On average, agents use 6.8 hours, 102 million tokens, and $43 to complete each task, with the longest attempts consuming 24 hours, 1.07 billion tokens, and $525. Notably, the highest-scoring agents used fewer tokens and were cheaper than less performant options. These results reveal that LLMs vary substantially in their ability to (1) execute multiple sequential analysis steps coherently, (2) manage large quantities of raw data, and (3) work across scientific domains.

Zane Koch, A. Wassie, Javier Valdes-Aleman et al. · 0 citations
Open access Aug 2026

A Persistent Fleet of AI Scientists Exhibits Cooperative and Autopoietic Behavior

A persistent fleet of cooperative AI scientist agents that operated continuously for nearly six months is described, term this bounded pattern AI Autopoietic Behavior due to the recurring operational improvement mediated by internal feedback and retained through institutional records with high confidence.

Milit S. Patel, W. Wierson, S. C. Ekker · 0 citations