Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality. Natural hazards make this work consequential because incomplete evidence can change estimates of severity, exposure, and mechanism. We introduce EarthVerse, a benchmark that evaluates sci...
Zhiqing Cui, Xinxiang Yin, Yihong Tang et al.· 0 citations
Evaluated on diverse mathematical, algorithmic, and systems optimization tasks, CORAL sets new state-of-the-art results on 10 tasks, achieving 3-10 times higher improvement rates with far fewer evaluations than fixed evolutionary search baselines across tasks.
Ao Qu, Handi Zheng, Zi-Jian Zhou et al.· arXiv.org· 38 citations· ⚡7
WorldCupArena is presented, a dynamic benchmark for language models and deep-research agents that can be reused for future leagues and cups, and shows only small gains in result and exact-score accuracy, but a clearer gain in Scoreline.
Zhaokai Wang, T. Gui, Jiayuan Rao et al.· arXiv.org· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.