Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models
Evaluating the capability and efficiency of LLMs from the DeepSeek-R1-Distill model family across four classes of arithmetic and algorithmic reasoning problems reveals potential limitations of naive scaling as a strategy for developing more capable AI systems.