Large Language Models are increasingly deployed as information intermediaries, yet measuring their political behavior remains fragile because questionnaire results mix model dispositions with measurement artifacts and response-elicitation biases. We introduce a robust Political Compass Test evaluation framework that samples 300 configurations across an eight-dimensional perturbation space varying language, framing, instructions, answer format, option order, and persona wording. We evaluate eight Gemma 3 and Qwen 3 models across 14 languages and three quantization levels, obtaining design-averaged political coordinates with quantified uncertainty. Most models lean Libertarian-Left on average, but instruction phrasing, language, and answer format significantly affect recovered coordinates. Cross-lingual differences primarily reflect coordinate drift rather than distinct cultural reasoning. Reverse-engineering the test also exposes axis-weighting imbalances and the collapse of degenerate responses toward the center, so near-origin estimates for the smallest models can reflect weak signal rather than centrism. Free-text reasoning and chat-then-classify elicitation alter recovered coordinates, and larger models show clearer persona separation, with a specific failure of the Authoritarian-Left persona to move most models in the intended social direction. In downstream tasks, persona effects are modest relative to model size and target group for hate-speech detection, while base and centrist prompts give the highest agreement for topic-level sentiment. Political role prompting therefore has measurable but task- and dataset-specific downstream effects.
Luka Debevc, Nishan Chatterjee, Antoine Doucet et al.· 0 citations
General-purpose Large Language Models (LLMs) like Llama, GPT, and Mistral struggle with domain-specific challenges in food and nutrition, where data is fragmented, heterogeneous, and semantically complex. While fine-tuned LLMs have shown success in healthcare and life sciences, similar progress in food domains has been limited, largely due to the lack of high-quality, task-specific datasets. We present FoodBench, a curated benchmark dataset of question–answer pairs designed for training and evaluating LLMs in food and nutrition. It spans key tasks such as nutrient estimation, food traffic-light classification, synonym linking, cooking measurement conversion, and food named-entity recognition and linking. FoodBench enables robust performance evaluation across zero-, one-, and few-shot settings, laying the groundwork for trustworthy, domain-adapted language models. This resource supports advances in personalized nutrition, dietary assessment, and food system innovation. Evaluation of four general-purpose LLMs (Llama 3, Mistral, Gemma, Gemini) on FoodBench tasks shows limited performance across nutrient estimation, traffic-light classification, and food interoperability, even with few-shot prompting. These results highlight the need for domain-specialized LLMs fine-tuned on food data, while establishing FoodBench as a benchmark not only for assessing general-purpose models but also for guiding and evaluating fine-tuning efforts.
T. Eftimov, Ana Gjorgjevikj, Matej Martinc et al.· Scientific Data· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.