Evaluation of large language model performance in translating dairy-related content
Abstract
Artificial intelligence’s ability to translate dairy-related texts from English to Spanish has not been well described in agriculture. This study aimed to determine the accuracy and comprehensibility of dairy-related translated text generated with ChatGPT (GPT-4o) and to determine the text’s appropriateness for a dairy employee. Eight dairy udder health and stockmanship-related English texts were gathered from extension and university websites for translation. Four texts were summaries of 260 words or less and the other four texts were procedure lists with 8 steps. The texts were translated with a large language model (ChatGPT-4o) and displayed in a side-by-side presentation. The presentations were sent to 23 reviewers with a rubric for evaluation of accuracy, comprehensibility, and perception of the translation. Additionally, the reviewers were asked to evaluate if the translations were suitable to be shared with dairy employees. Overall, reviewers “Strongly agreed” or “Agreed” that the texts were accurately translated, including technical terms, and that the main idea of the text was correctly reproduced. They also “Strongly agreed” or “Agreed” that the translation was easy to understand, flowed naturally, was clear and coherent, and there was consistent translation of phrases that were repeated. Additionally, reviewers believed that there were “None” or “Few” instances where the meaning of the text is unclear due to the translation. Finally, the reviewers mostly agreed that they would present the translations to dairy employees. Thus, ChatGPT (GPT-4o) can accurately and comprehensibly translate dairy-related text. However, it is important to still have oversight and review the output.