From Atomic to Agentic: Towards Interpretable Evaluation of LLMs'Agentic Mathematical Capabilities
Experiments reveal that models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles, demonstrating that process-level evaluation is crucial for interpreting the true potential of LLMs and guiding the development of next-generation mathematical agents.