Skip to content
Review Open access

From Prompt to Pipeline: A Comparative Evaluation of Large Language Model Coding Agents for Reproducible Bioinformatics Pipeline Construction

Sep 2026 · bioRxiv · 0 citations · 6 references
Biology

TL;DR

The results suggest that current coding agents can accelerate scaffolding, documentation, and routine implementation, but they do not eliminate the need for expert review in bioinformatics workflow construction.

Abstract

Agentic coding systems are increasingly presented as a way to reduce the engineering burden of scientific software development. Bioinformatics is a strong test case for this claim because useful pipelines must combine domain-specific analysis choices, command-line software, sample metadata, workflow orchestration, container or HPC execution, and interpretable quality-control reporting. We evaluated three agentic systems - Biomni, Claude Code, and Codex - on the same task: constructing a Nextflow DSL2 pipeline for paired-end CUT&Tag data that included read QC, trimming, alignment, filtering, duplicate removal, signal track generation, per-sample and group-level peak calling, control-aware group merging, annotation, FRiP calculation, deepTools visualizations, and final MultiQC reporting. Each system received the same detailed CRAFT-style prompt and was assessed against a hand-coded reference pipeline developed by the authors. All three systems produced pipeline implementations that appeared plausible at the level of documentation and file structure, but none fully satisfied the requested analysis. The most consequential failure was shared: the agent-generated pipelines performed some form of group-level merging but did not produce the requested merged-group reporting outputs. Sample-level MultiQC reports also disagreed with the reference report. Codex was closest to the reference for primary mapped-read counts, although its total-read accounting and report structure still differed. Claude Code produced the broadest final report, but its mapping summary mixed stages and therefore could not be treated as numerically correct. Biomni produced the strongest subjective documentation, but its read-count agreement with the reference report was poor and several failures required substantial Nextflow expertise to diagnose. These results suggest that current coding agents can accelerate scaffolding, documentation, and routine implementation, but they do not eliminate the need for expert review in bioinformatics workflow construction. For complex sequencing workflows, prompts must specify not only the biological intent, but also the exact stage semantics, acceptance tests, metadata contracts, expected report sections, resource propagation rules, and failure criteria needed to distinguish a plausible pipeline from a correct one. Author summary Modern AI coding agents can write large amounts of software from natural language instructions. This study asks whether that ability is enough to build a real bioinformatics pipeline for CUT&Tag sequencing data. The answer from this evaluation is mixed. The agents produced impressive-looking pipelines and documentation, but the outputs did not fully match a hand-coded reference workflow. Most importantly, the pipelines did not generate the requested merged-group reports, which were central to the biological use case. Some errors were simple programming issues, but others required a working knowledge of Nextflow, sequencing QC, and chromatin profiling workflows. The practical conclusion is not that these tools are useless. Rather, they are best viewed as accelerators for experienced users. A biology student with little programming or workflow-management background would still have difficulty detecting and fixing the most important errors.

Read PDF

Similar papers

Review Open access Sep 2026

Software engineering for reproducible pipeline development in bioinformatics

Reproducibility in bioinformatics remains challenging despite the availability of workflow management systems and mature computational infrastructures. This work presents a software-engineering perspective for developing reproducible bioinformatics pipelines, with emphasis on pipeline-specific code. We reinterpret the...

D. Pérez-Rodríguez, Alba Nogueira-Rodríguez, Jorge Vieira et al. · 0 citations
Open access Aug 2026

REAPER: a project-centric workflow layer for comparative repeatome analysis

This work presents REAPER (Repeatome Extended Analysis Pipeline—Execution and Reporting), a project-centric workflow layer that couples a modular Snakemake pipeline with a Python project manager to enforce a stable on-disk layout and configuration-driven execution for single-sample and comparative repeatome analyses.

D. Ulyanov, A. I. Yurkina, Viktoria Voronezhskaya et al. · 0 citations
Book Open access Aug 2026

Benchmarking LLM Agents on Real-World Biological Database Curation for Data-Driven Scientific Discovery

BioDataLab evaluates the capability of autonomous agents to transform raw, heterogeneous biological resources into structured, analysis-ready databases, and underscores that while LLMs are proficient in downstream reasoning, autonomous upstream curation remains a formidable frontier.

Jiaxian Yan, Xi Fang, Jintao Zhu et al. · 0 citations
Open access Aug 2026

Automating scientific annotations for open transcriptomic profiles via multi-stage agents

GEOMeta provides a scalable resource and reproducible framework for metadata curation in the Gene Expression Omnibus, and benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings.

Xiaodan Zhang, S. Paithankar, Jing Pu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Bioinfoysis Technical Report

A multi-agent harness that represents each request as a persistent, artifact-grounded analysis run, and demonstrates that reliable bioinformatics automation depends not only on model capability, but also on the harness that governs planning, execution, memory, and evidence flow.

Qi Shao, Xin Zhang, Zhou-Yang Yuan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.