Preprint
Aug 2026
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models
A large-scale assessment of the effectiveness and robustness of these automated pipelines is conducted by evaluating five widely used benchmark suites across 26 open-source SLMs under a unified judging rubric, which reveals a capability-safety confound that mixes model capability with apparent safety.
Nyamtulla Shaik, Fengjun Li, Bo Luo
· 0 citations