Preprint
Aug 2026
Lost in Aggregation: How Benchmarks Overlook Irreplaceable Model Strengths
This work argues that benchmark evaluation should also consider the data-centric peak performance frontier, defined by the best statistically supported performance achieved on each dataset, and finds that common aggregation metrics are highly correlated and largely measure consistency and avoiding failures, while being much less aligned with dataset-level irreplaceability.
Andrej Tschalzev, Stefan Lüdtke, Heiner Stuckenschmidt et al.
· 0 citations