Efficiently summarizing dietary records at scale remains a persistent bottleneck in nutritional epidemiology. We present FoodScribe, which translates free-text meal descriptions into quantitative nutrient profiles by combining ingredient parsing with nutrient retrieval by querying the USDA FoodData Central (FDC) database. Benchmarked using three LLM providers using Nutribench dataset, FoodScribe completed annotation of 3,807 meal descriptions in 2.5 hours, a task otherwise requiring substantial manual effort from trained nutritionists. FoodScribe achieved accuracy across macronutrient estimation (F1=0.79-0.89), with models performing better for protein than fat estimation. Application to a Mediterranean diet intervention cohort indicated dietary shifts consistent with the intervention pattern based on model-derived estimates. Integration with metabolomics data suggested that fiber and vegetable intake were positively associated with a fecal metabolite cluster.
Harsha Gouda, M. Sala-Climent, Julius Agongo et al.· medRxiv· 0 citations
Integrating multi-omics data is essential for microbiome research, as microbial communities are shaped by and respond to interdependent processes, including taxonomic composition, metabolite production and utilization, and gene expression. However, accurately capturing ecosystem-wide patterns across these modalities is statistically challenging due to differences in scale, sparsity, and compositionality. While a growing number of multi-omics methods have emerged, they differ in their mathematical objectives and modeling assumptions, which in turn shape how biological patterns are represented and interpreted. This underscores the need for tools that explicitly account for the statistical properties of microbial ecosystems. Here, we present Joint Robust Principal Component Analysis (Joint-RPCA), a method designed with these statistical properties in mind and broadly applicable to multi-omics settings with similar challenges. Built on the OptSpace matrix completion framework, Joint-RPCA assumes an underlying shared low-rank structured component across modalities to identify shared variation and cross-modal associations from matched samples. Within this setting and under these statistical assumptions, Joint-RPCA showed stronger performance than the benchmarked general-purpose methods in phenotype separation and feature association tasks, achieving up to sixfold improvement in classification accuracy and over 100-fold faster runtimes. Applied to real-world datasets, including the Integrative Human Microbiome Project (iHMP), mammalian gut microbiomes, and decomposition studies, Joint-RPCA reveals replicable and interpretable multi-omic patterns, offering a scalable and domain-aware solution for systems-level microbiome analysis. Joint-RPCA is available in both Python ( https://github.com/biocore/gemelli ) and R ( https://bioconductor.org/packages/mia ).
Bianca Cordazzo Vargas, C. Martino, A. Dilmore et al.· Molecular Systems Biology· 1 citation