A Japanese-Dietitian Prompt Systematically Shifts Portion Estimates in LLM-Based Nutrient Estimation from Food Images: A Multi-Dataset, Multi-Model Study
Background/Objectives: In food-image nutrient estimation with vision-language models (VLMs), portion size is a dominant source of error, yet how a prompt shifts the estimated amount—and in which direction—remains uncharacterized. We tested whether a Japanese-dietitian prompt acts as a systematic, directional influence on a model’s quantity estimates and what an explicit magnitude instruction does by comparison. Methods: Six VLMs from three vendors were evaluated on two datasets with contrasting portion regimes—NutriImage (Japanese cafeteria dishes; dietitian-calculated ground truth) and SNAPMe (US meal photographs)—under persona conditions (none, Japanese-dietitian, US-dietitian) crossed with two portion-specification levels, with five additional prompt-control conditions, image-level paired statistics with bootstrap confidence intervals, interaction tests, equivalence tests, and a repeated-call variability analysis. Results: The Japanese-dietitian prompt lowered predicted energy in all six models and both datasets (persona main effect p < 10−94), approximately preserving predicted macronutrient composition in relative terms; the US prompt produced only small, sign-inconsistent changes. The effect was not reproduced by an explicit “assume smaller portions” instruction, which shifted estimates further but far less consistently, whereas an “assume larger portions” instruction was followed almost uniformly by four of the six models, with both OpenAI models largely insensitive to explicit magnitude instructions in either direction. Accuracy consequences were dataset-dependent: error decreased on the small-portion dataset (up to ~20 MedAPE points) and was statistically equivalent (±5-point margin) on the larger-portion dataset in 11 of 12 model × portion cells. All principal effects were confirmed in a unified analysis that randomizes model identity over a pooled image set (n = 2159 independent images; Holm-corrected; robust across 1000 random re-assignments). Conclusions: A Japanese-dietitian prompt acts as a consistent downward influence on VLM quantity estimates that is distinct from explicit downscaling instructions; its accuracy value is domain-specific and requires validation on the target domain before any practical use.