Preprint
Jul 2026
SLBench: Evaluating How LLM Agents Follow Logical Relations in Skills
This work introduces SkillLogic, a framework for analyzing logical relations in skill files and constructing executable tests from them, and establishes logical-relation following as a distinct reliability challenge for skill-guided agents.
Xuan Chen, Chengpeng Wang, Lu Yan et al.
· 0 citations