Skip to content
Review

Evaluation of Energy and Nutrient Estimates from Large Language Models Using Text-Based Queries.

Jul 2026 · Journal of NutriLife · Vol 156, pp. 101712 · 0 citations · 48 references
Medicine

TL;DR

LLMs show promise for estimating energy and macronutrients, however, performance for micronutrients requires further improvement and may affect overall dietary assessment.

Abstract

Background

Large language models (LLMs) have emerged as promising tools for estimating energy and nutrient values, yet most existing evaluations focus on image-based queries rather than text. Few studies compare LLM estimates with reference databases commonly used in nutrition research.

Objective

To examine agreement between LLM estimates and a research food composition database for energy and nutrients, and to determine if agreement varies by food group.

Methods

We conducted a cross-sectional analysis of energy and nutrient estimates for frequently consumed food items in the United States (US). Food items were entered as text prompts into four LLMs (ChatGPT 5.2, Claude Opus 4.5, Gemini 3 Pro Preview, and Llama 4 Maverick), which provided energy and nutrient estimates. Corresponding foods were matched to the Nutrition Coordinating Center (NCC) Food and Nutrient Database, and agreement between LLM and database values was assessed using intraclass correlation coefficients (ICCs) and Bland-Altman analyses. Agreement was also evaluated within the three most frequently consumed food groups.

Results

Agreement was high for energy and macronutrients for all LLMs. Variability was observed for several micronutrients, particularly vitamin D, folate, and iron. Claude Opus 4.5 showed consistently high agreement, with no nutrients classified as poor. Other LLMs exhibited poor agreement for at least one micronutrient. Certain food categories, including condiments and mixed dishes, contributed disproportionately to variability. However, agreement remained high within the most frequently consumed broader food groups.

Conclusions

LLMs show promise for estimating energy and macronutrients. However, performance for micronutrients requires further improvement and may affect overall dietary assessment.

View source

Similar papers

Open access Jul 2026

FoodScribe: an open-source semantic framework for nutrient estimation from free-text dietary records

Efficiently summarizing dietary records at scale remains a persistent bottleneck in nutritional epidemiology. We present FoodScribe, which translates free-text meal descriptions into quantitative nutrient profiles by combining ingredient parsing with nutrient retrieval by querying the USDA FoodData Central (FDC) database. Benchmarked using three LLM providers using Nutribench dataset, FoodScribe completed annotation of 3,807 meal descriptions in 2.5 hours, a task otherwise requiring substantial manual effort from trained nutritionists. FoodScribe achieved accuracy across macronutrient estimation (F1=0.79-0.89), with models performing better for protein than fat estimation. Application to a Mediterranean diet intervention cohort indicated dietary shifts consistent with the intervention pattern based on model-derived estimates. Integration with metabolomics data suggested that fiber and vegetable intake were positively associated with a fecal metabolite cluster.

Harsha Gouda, M. Sala-Climent, Julius Agongo et al. · 0 citations
Review Open access Jul 2026

Vision-Language Models for Image-Based Dietary Assessment: A Benchmark of Accuracy, Cost, and Prompt Strategies Across Ten Models

Background Dietary assessment is the cornerstone of clinical management and research studies evaluating diet and health. Traditional methods such as food diaries and 24-hour recalls can be burdensome, prone to recall bias, and difficult to adhere to. Image-based dietary assessment using vision-language models (VLMs) offers a potential solution. Objective Our goal was to benchmark state-of-the-art VLMs for automated food recognition, weight estimation, and calorie estimation using Google’s Nutrition5k dataset. Methods We evaluated 3,229 food images using ten approaches: proprietary VLMs (Gemini 2.0 Flash, 2.5 Flash, 3.0 Flash, and 3.1 Flash Lite; GPT- 4o, GPT-4o-mini, and GPT-5 Mini; and Claude Haiku 4.5), an open-source VLM (Qwen2-VL-7B), and a commercial food recognition API (FatSecret). We assessed calorie and weight estimation using Lin’s Concordance Correlation Coefficient (CCC) and component detection using Jaccard similarity. Results Gemini 3.0 Flash achieved the best calorie estimation (CCC 0.767, MAE 80.7 kcal), while Gemini 3.1 Flash Lite offered very comparable accuracy (CCC 0.754) with the highest ingredient recognition (Jaccard 0.655) at the lowest cost among top-performing models ($0.59/1K images). Among earlier-generation models, Gemini 2.0 Flash remained competitive (CCC 0.742, Jaccard 0.621) at a fraction of the cost ($0.10/1K images). A human validation study in which four annotators reviewed 440 images revealed systematic omissions in the original Nutrition5k labels. After correction, the extrapolated ingredient-overlap score for Gemini 2.0 Flash increased from 0.62 to an estimated 0.82, suggesting that raw Jaccard scores substantially underestimate true model performance. Conclusions Current VLMs can perform automated dietary assessment with reasonable accuracy from single overhead photographs. Our results inform model selection for dietary assessment applications and highlight remaining challenges in calorie estimation and component detection for complex, multi-item meals.

Sam Sterling, Lauren T. Berube, Andrea J. Glenn et al. · 0 citations
Aug 2026

Evaluating LLM Accuracy in Predicting Peruvian Meal Nutrition.

BACKGROUND Artificial intelligence applications have been developed to predict the nutrient content of meals. However, none have been evaluated in the context of Peruvian cuisine, characterized by diverse ingredients and recipes. We assessed whether large language models (LLMs) could predict the nutritional content of Peruvian meals. METHODS Using a dataset of 510 unique lunch images extracted from a Peruvian cookbook, we compared nutrient values from recipe data against predictions generated by LLMs (Gemma-3 4B, 12B, and 27B). The LLMs were given the meal name and a photograph and prompted to produce narrative descriptions of the meal. Using the descriptions, the same LLMs were prompted to estimate six nutrients: energy (kcal/serving), protein (g/serving), carbohydrates (g/serving), iron (mg/serving), vitamin A (μg/serving), and zinc (mg/serving). Agreement proportions and errors metrics were calculated against the values from the recipe book. RESULTS The 27B LLM achieved the highest agreement proportions across most nutrients-calories (45%), carbohydrates (31%), iron (15%), vitamin A (19%), and zinc (31%)-while the 12B model performed best for protein (70% agreement). The 27B model yielded the lowest mean absolute error (MAE) for calories (108 kcal), carbohydrates (26 g), iron (4 mg), and zinc (1 mg). The 12B LLM had the lowest MAE for protein (6 g) and vitamin A (667 μg). The 4B LLM showed the poorest performance across metrics. CONCLUSIONS LLMs can generate estimates of nutrient content from narrative descriptions of Peruvian meals, but current performance levels fall short of the precision required for clinical deployment or commercial consumer-facing applications.

R. M. Carrillo-Larco, Mariano Gallo Ruelas, Mika Matsuzaki et al. · 0 citations
Review Open access Aug 2026

Evaluation of large language model performance in translating dairy-related content

Artificial intelligence’s ability to translate dairy-related texts from English to Spanish has not been well described in agriculture. This study aimed to determine the accuracy and comprehensibility of dairy-related translated text generated with ChatGPT (GPT-4o) and to determine the text’s appropriateness for a dairy employee. Eight dairy udder health and stockmanship-related English texts were gathered from extension and university websites for translation. Four texts were summaries of 260 words or less and the other four texts were procedure lists with 8 steps. The texts were translated with a large language model (ChatGPT-4o) and displayed in a side-by-side presentation. The presentations were sent to 23 reviewers with a rubric for evaluation of accuracy, comprehensibility, and perception of the translation. Additionally, the reviewers were asked to evaluate if the translations were suitable to be shared with dairy employees. Overall, reviewers “Strongly agreed” or “Agreed” that the texts were accurately translated, including technical terms, and that the main idea of the text was correctly reproduced. They also “Strongly agreed” or “Agreed” that the translation was easy to understand, flowed naturally, was clear and coherent, and there was consistent translation of phrases that were repeated. Additionally, reviewers believed that there were “None” or “Few” instances where the meaning of the text is unclear due to the translation. Finally, the reviewers mostly agreed that they would present the translations to dairy employees. Thus, ChatGPT (GPT-4o) can accurately and comprehensibly translate dairy-related text. However, it is important to still have oversight and review the output.

Anay D. Ravelo, Daniela G. Carranza, Kaitlyn A. Lutz et al. · 0 citations
Open access Jul 2026

Laboratory-Analyzed Nutrient Values Diverge from Database Estimates in Research-Specific Diets (Notably for Protein, Iron, Sodium)

Nutrient databases are widely used to calculate the composition of research diets, yet their accuracy relative to laboratory measurements remain uncertain. We compared database-calculated values with laboratory-analyzed energy, protein, fat, and trace mineral concentrations in two controlled 5-day dietary patterns—the Dietary Guidelines for Americans Healthy Eating Pattern (DGA-HEP) and a typical American eating pattern (TAEP)—prepared in independent triplicate and analyzed as technical replicates, using a descriptive approach with a >10% relative-difference threshold. Energy and fat generally agreed between laboratory analyses and database estimates, whereas protein was consistently higher by laboratory analysis (DGA-HEP: +13 ± 2%; TAEP: +25 ± 5%). Trace minerals showed nutrient-specific discrepancies, with sodium systematically higher by laboratory analysis (~+14% to +40% across days) and iron systematically lower (~−15% to −58% across days) for both patterns; other minerals varied by nutrient and day. These results indicate that nutrient databases may over- or underestimate key nutrients depending on category, brand, formulation, and preparation. When precise composition is essential for study design or interpretation, targeted laboratory validation—particularly for protein, sodium, and iron—may be warranted.

Emma Patzer, Maggie Dervis, A. Scheett et al. · 0 citations
Open access 2026

Assessing Hospital Patient Nutrient Intake with an AI-Powered Food Recognition System – A Feasibility Study of the FlavoriaFlex solution

Adequate dietary intake is essential for positive clinical outcomes of hospitalized patients, yet monitoring food intake is labor-intensive and often subjective. AI-based food recognition could automate monitoring and assessment, but evidence in real-world hospital settings is limited. This study evaluated an AI-powered food recognition system, FlavoriaFlex, to assess its detection performance, deployment feasibility, and acceptability among dietitians. Previously validated in restaurant (F1 0.75, weight MAE 23.6 g, energy MAE 235 kcal), the system was deployed in a hospital ward for six days. A total of 133 meals were recorded; 102 had paired leftover images (235 total images). Manual annotation of 483 food segments provided ground truth for evaluating food recognition and menu mapping. Semi-structured interviews with dietitians assessed usability, perceived benefits, and clinical value. FlavoriaFlex enabled automatic estimation of item- and meal-level consumption, including weights and energy- and macronutrient contents. Overall food recognition accuracy was 94% (F1 0.76), remaining high for served meals (96.5%, F1 0.85) and robust for visually complex leftovers (89.5%, F1 0.71). Unknown/non-food segments were minimal (2.4% of leftovers; 0.27% of weight). A web dashboard delivered real-time visualizations, including energy and nutrient intake. Dietitians reported reduced cognitive burden, more objective assessment, and improved observability into patient dietary intake, while emphasizing the need for further validation and integration for clinical use. These findings demonstrate that FlavoriaFlex could be integrated into hospital workflows to provide accurate, clinically meaningful intake estimates, with AI-assisted food recognition offering an efficient, reliable approach to improving nutritional monitoring at scale.

Rehan Khalil, Sanna Koskimäki, Hanna Lähde et al. · 0 citations