Preprint
Jul 2026
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
OmniaBench provides a broad and diagnostic benchmark for characterizing the capability boundaries of general agents across diverse scenarios with explicit state spaces, and introduces a ten-dimensional capability taxonomy and eight compositional atomic difficulty factors to support fine-grained evaluation and analysis.
Chengyu Shen, Yujie Fu, Gangtao Xin et al.
· 0 citations