Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks are typically short-horizon, requiring only a small number of retrieval or inference steps, and provid...