Skip to content
Review Open access

Agentic genomics: From pipeline automation to autonomous validation

Jul 2026 · Cell Genomics · Vol 6, pp. 101305 · 0 citations · 38 references
Medicine

TL;DR

It is argued that agentic genomics shifts the bottleneck in computational biology from pipeline construction to validation and proposes a tiered validation framework spanning research-grade, benchmarked, and clinical-grade analyses and argues that equity-aware design must be a systems requirement rather than an optional aspiration.

Abstract

Summary Genomics has entered a phase in which AI agents can autonomously discover, configure, execute, and chain bioinformatics operations from natural-language instructions. We term this paradigm “agentic genomics”: the delegation of multi-step genomic analyses to autonomous software agents that select tools, manage dependencies, and adapt execution in response to intermediate results, mediated by large language models (LLMs) and constrained by domain-specific skill libraries. We argue that agentic genomics shifts the bottleneck in computational biology from pipeline construction to validation. We examine emerging systems, including CellAtria, AutoBA, Bio-Copilot, and ClawBio, and assess their divergent architectures. We propose a tiered validation framework spanning research-grade, benchmarked, and clinical-grade analyses and argue that equity-aware design must be a systems requirement rather than an optional aspiration. We identify the infrastructure needed to make agentic genomics trustworthy.

Read PDF

Similar papers

Review Jul 2026

Evaluating Agentic Bioinformatics through Function, Evidence, and Validation

It is argued that agentic bioinformatics should be assessed through workflow correctness rather than final-answer correctness alone, and the Function--Evidence--Validation (FEV) framework is introduced, which separates demonstrated workflow operations, traceable support for actions and claims, and use-case-specific validation.

Phuc Pham, Truong-Son Hy · 0 citations
Open access Jul 2026

Agentic AI integrated with scientific knowledge: laboratory validation in systems biology.

Automation is transforming scientific discovery by enabling systematic exploration of complex hypotheses. Large language models (LLMs) perform well across diverse tasks and promise to accelerate research, but often struggle with logical structures. Here, we present a framework for biological discovery integrating LLM-based agents with laboratory automation, guided by logical scaffolds incorporating symbolic relational learning, structured vocabularies and experimental constraints. This integration improves coherence and reliability in automated workflows. We couple this AI-driven approach to automated cell-culture and metabolomics platforms, enabling integrated hypothesis validation and refinement, yielding a flexible discovery system. The system identified novel interactions in Saccharomyces cerevisiae, including glutamate-induced growth inhibition in spermine-treated cells and aminoadipate's partial rescue of formic-acid stress. All hypotheses, experiments and data are captured in a graph database employing controlled vocabularies. Existing ontologies are extended, and a novel representation of scientific hypotheses is presented using description logics. This work demonstrates the potential for a reliable machine-driven discovery process in systems biology.

Daniel Brunnsåker, Alexander H. Gower, Prajakta Naval et al. · 2 citations
Review Open access Jul 2026

From Executor to Orchestrator: The Pharmacology Scientist in the Age of Agentic AI

Drug development productivity has not improved despite five decades of computational advancement, with the probability that a compound entering Phase I achieving regulatory approval remaining near 10%. Each automation wave increased throughput while leaving the interpretive bottleneck intact; scientists continued to formulate questions, evaluate outputs, and make advancement decisions regardless of how fast data accumulated upstream. Agentic AI systems capable of reasoning, planning, and executing multi‐step analyses without continuous human instruction represent the first class of ubiquitous computational tools with the architectural potential to compress this bottleneck, but the implications extend beyond efficiency. As computational systems begin to perform interpretation, execution, and evaluation steps that previously required human judgment, the scientist's contribution shifts from conducting analyses to specifying objectives precise enough for autonomous execution and evaluating recommendations that may be difficult to verify independently. Whether this shift improves aggregate productivity depends on whether autonomous systems address the fundamental causes of clinical failure, including insufficient efficacy, inadequate safety prediction, and poor preclinical translation, rather than merely accelerating the analytical work surrounding them. This review examines what clinical pharmacology scientists must become as these systems enter routine practice. The competencies required for effective orchestration differ from those emphasized in traditional pharmaceutical training, and existing governance structures do not address the failure modes that accompany delegation of scientific judgment to autonomous systems. Whether this shift improves productivity or introduces new failure modes depends on governance and training investments that the field has not yet made.

Michael G. McCoy, Matthew McCoy · 0 citations
Open access Aug 2026

Multi-Agent Readiness Scoring Methodology in Bioinformatics Domain

The emergence of Large Language Models (LLMs) has significantly advanced computational biology, yet their integration into autonomous, multi-agent systems (MASs) and clinical workflows remains challenging due to systemic architectural fragmentation. To quantify the operational readiness and regulatory compliance of bioinformatics LLMs, we developed the Multi-Agent Readiness Score (MARS), a standardized evaluation framework assessing models across four structural dimensions: Governance & Accessibility, Biological Competence, Technical Maturity, and Agentic Orchestration. The framework incorporates compliance criteria from the EU AI Act, HL7 FHIR, HL7 CDA, and MyHealth@EU standards. To empirically validate this domain-agnostic methodology, we applied it to a highly mature subset of the field: a diverse cohort of 43 prominent genomic LLMs. Our assessment revealed a severe, industry-wide readiness gap: the majority of models fell into “Not Suitable” or “Research Prototype” tiers, lacking essential technical interfaces, structured communication schemas, and provenance tracking. Furthermore, the data demonstrated a ’competence-readiness gap’, where models scale in biological predictive competence without corresponding improvements in engineering utility. The primary barrier to scalable bioinformatics AI is no longer biological competence, but operational and architectural incompatibility. By quantifying integration friction, MARS provides a crucial, reproducible metric to audit model maturity, guide system architecture, and ensure future models are structurally prepared for the rigorous regulatory demands of precision medicine workflows.

Blagojche Gjorgjioski, Djansel Bukovec, Ivana Vichentijevikj et al. · 0 citations
Preprint Aug 2026

DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data

Clinical machine learning (ML) has the potential to support high-stakes medical decision-making, but reliable deployment is often constrained by scarce, heterogeneous, and temporal complexity. Developing effective ML pipelines for such data remains time-consuming and error-prone, while existing automated machine learning (AutoML) systems only partially address this challenge because they largely rely on brute-force search over predefined spaces and lack explicit reasoning and memory. We therefore reformulate AutoML for small clinical data from exhaustive search to reasoning-driven refinement. We propose DoctorAgents, an agentic AI framework that autonomously constructs and optimizes end-to-end ML pipelines through specialized large language model (LLM) agents for generation, validation, and refinement. DoctorAgents backpropagates natural-language feedback through textual gradient descent to perform targeted updates without exhaustive search. Experiments across diverse clinical tasks show that DoctorAgents consistently outperforms established AutoML baselines while producing more interpretable task-specific representations.

Ruilin Wang, Bozhong Wang, Elizabeth Kourbatski et al. · 1 citation
Preprint Aug 2026

CAi Copilot: Reducing Operational Workload in Molecular Design through Intent-Driven Agentic Workflows

Early-stage molecular design is an iterative process, not just a task of generating molecules. Researchers turn broad goals into design strategies, refine candidates, assess many properties, and gather evidence before synthesis and tests. AI methods can generate molecules, optimize several goals, predict properties, dock compounds, and account for synthesis. Yet these functions are spread across specialized tools. Experts must still coordinate each step, judge interim results, and integrate evidence. The central challenge is thus to turn research intent into adaptive, traceable runs grounded in scientific tools. We cast this challenge as intent-to-evidence molecular design workflow execution and present CAi Copilot, an expert-oriented agent with three linked layers. The Research Interface Layer turns intent into an executable plan. The Agent Reasoning Layer uses interim results to guide each run. The Execution Substrate supplies molecular tools, metrics, reusable utilities, and backend services. Across 45 tasks, CAi achieves the strongest overall performance, with an outcome score of 84.59, exceeding the next-best result by 18.07 points. Additional benchmarks test how CAi coordinates generation, screening, and multi-criteria evaluation, while exposing limits in long-horizon execution. These results show that CAi turns broad molecular-design intent into transparent, traceable workflows that connect interim decisions to candidate-level evidence.

Zhuo-Lei Wang, Jiangyu Chen, Yingjun Shang et al. · 0 citations