Preprint
Aug 2026
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning
Reinforcement fine-tuning strategies for adapting Qwen2.5-3B-Base to graduate-level signal mathematical problems from WirelessMATHBench-XL are investigated and whether GSPO or GMPO offer advantages in stability or accuracy over GRPO for signal reasoning tasks is assessed.
Guozheng Sun
· 0 citations