Oct 2026· PLoS Computational Biology· Vol 22 10, pp.
e1014822
· 0 citations· 36 references
Medicine
Abstract
The biomedical literature contains a vast collection of omics studies, yet most published data remain functionally inaccessible for computational reuse. When raw data are deposited in public repositories, essential information for reproducing reported results is dispersed across main text, supplementary files, and code repositories, and in the rarer cases where intermediate data (e.g., protein abundance files) are shared, their location is irregular. Here we present an agentic framework for the agent-supported retrieval, reanalysis, and synthesis of published omics data. The system employs large language model (LLM) agents with access to tools for fetching omics studies, extracting article metadata, identifying and downloading published data, executing containerized quantification pipelines, and synthesizing results across studies. Applied at corpus scale, the pipeline catalogued dataset references across thousands of PubMed Central articles; we report these as descriptive system outputs rather than as a validated measure of extraction accuracy. Using model context protocol (MCP) servers to expose containerized analysis tools, the agents retrieved and re-quantified data in five end-to-end reanalyses spanning data-dependent and data-independent proteomics and bulk RNA-seq. All five reanalyses completed, each with documented human guidance and workflow accommodations, and reproduced the authors' deposited abundances with high per-sample correlation (0.85-0.997) and strongly concordant differentially expressed features (fold-change Spearman 0.88-0.91), with no direction reversals among features called differentially expressed in both analyses; residual differences in significant-feature lists were attributable to threshold placement, tool-version, and preprocessing differences rather than to the underlying quantities. We further demonstrate that agents can identify semantically similar studies, judge data compatibility, and synthesize findings across studies, including a random-effects meta-analysis that recovered consistent protein regulation in liver fibrosis. Rather than a validated benchmark of literature-wide performance, this work is a feasibility demonstration together with an auditable, reusable toolset, establishing a foundation for prospective evaluation of automated omics-data reuse.
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.
P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al.· IEEE International Conferenc...· 110 citations· ⚡7
The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· International Conference on...· 84 citations· ⚡6