Improving Mathematical Reasoning in Small Language Models via Curriculum Learning and Iterative Execution Feedback
These extensions focus on two approaches: curriculum learning and inference-time iterative feedback, where a Qwen-2.5-7B-Instruct model is used as a critic to provide corrective advice on previous round’s incorrect answers, allowing the model to learn from its own mistakes and make revisions.
Jiayu Sui, X. Ai
· 0 citations