Skip to content
Open access

MetSEval-1k: A Comprehensive Benchmark for Evaluating Large Language Models in Meteorology

Aug 2026 · Applied Sciences · 0 citations · 20 references

Abstract

This paper proposes METMAP, a comprehensive 6D evaluation framework designed to assess large language models in meteorological applications. The framework encompasses six critical capabilities including meteorological knowledge comprehension, expert-level meteorological content summarization, multilingual translation of meteorological information, geospatial mapping context understanding, alignment with authoritative meteorological standards, and professional meteorological service communication. To operationalize this framework, the paper further introduces MetSEval-1k, a high-quality benchmark comprising 1083 expert-curated questions spanning operational and public-facing meteorological services. The benchmark integrates both objective multiple-choice items and subjective open-ended tasks to enable holistic model assessment. This paper conducts systematic evaluations of multiple state-of-the-art large language models using MetSEval-1k, revealing substantial performance disparities across the six dimensions. The results highlight the critical necessity for domain-specific adaptation and rigorous validation before deploying large language models in operational meteorological contexts. MetSEval-1k is released as a foundational benchmark to advance research and development of trustworthy, service-oriented artificial intelligence application in meteorology.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.