Skip to content

Author

J. Sáez-Rodríguez

We have 3 of 329 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Feb 2025

Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data

A fundamental design pattern in biomolecular studies is to assay the same set of samples (organisms, tissue biopsies, or individual cells) by multiple different ‘omics assays. Group Factor Analysis (GFA) and its adaptation to high-dimensional settings, Multi-Omics Factor Analysis (MOFA), are widely used as a first-line approach to analyse such data and are effective in detecting patterns of correlation, organize them into so-called latent factors, and identify common and assay-specific factors. However, in many applications a subset of the found factors just rediscovers already known covariates (e.g., disease subtypes, environmental covariates) while others may represent genuine novelty. Here, we present Semi-supervised Omics Factor Analysis (SOFA), a method that incorporates known covariates into the model upfront and focuses the factor discovery on novel sources of variation. We show SOFA’s effectiveness for discovering novel patterns by applying it to cancer, brain development and heart failure multi-omic data sets.

Tümay Capraz, Harald Vöhringer, Klaus Sebastian Augusto Kruger Serrano et al. · 2 citations
Preprint Jul 2026

EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast. Access to the literature is a crucial part of that need, and resources such as Europe PMC, with over 40M indexed records, are widely used to meet it. Yet these resources were not built for AI agents: they take keywords and complex syntax and return whole papers, so every agent must learn the syntax, issue several searches, and read full papers to find the evidence it needs. We introduce EMBL AI Librarian, a knowledge layer that upgrades the Europe PMC interface for AI agents: an agent asks in natural language and receives evidence that answers it. A single LLM orchestrates the whole knowledge retrieval process: it plans complementary subqueries executed by the live Europe PMC search engine, then reads the selected papers and locates the relevant evidence. We evaluate Librarian across four benchmarks: literature synthesis, claim verification, open-domain question answering, and downstream biology tasks such as protocol questions and sequence manipulation. On ScholarQABench, Librarian improves Citation F1 by more than $16$ points over strong recently published baselines. Used as the retrieval layer of an existing claim-verification pipeline, it increases agreement with expert consensus; and on the open-form LitQA2 benchmark, a GPT-5.4 agent scores about $8$ points higher when grounded in Librarian than with web search. Overall, our results show that equipping life-science agents with the Librarian knowledge layer improves performance across a range of tasks. We release our code publicly at https://github.com/petroni-lab/librarian

Luigi Sigillo, M. Silvestri, Francesco Tabaro et al. · 0 citations
Open access Jul 2026

Multi-modal data integration reveals functionally credible predictive biomarkers in ovarian cancer.

An integrated whole-genome and transcriptome workflow is developed to systematically distinguish functionally credible, predictive driver aberrations from non-functional alterations across all classes of genomic events in ovarian high-grade serous carcinoma.

T. Muranen, A. Hainari, D. Afenteva et al. · 0 citations