The Model Says Walk: Measuring whether LLMs Condition on Hidden Constraints
Standard accuracy flatters all ten models the authors evaluate and reorders their ranking, and prompting fixes that look effective largely vanish under paired scoring.
Yubo Li, Lu Zhang, Tian-Chong Jiang et al.
· 0 citations