Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions
An exam-style evaluation framework is introduced for studying the global budget allocation of reasoning language models when multiple problems share an end-to-end cost or latency constraint, in which a model must distribute one shared token budget across questions with different difficulty and point values.